Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

311–320 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#311
post #195

Earlier quoted context omitted.

It's not okay for a human to pirate, plagiarize, violate IP rights and laws, etc. But I disagree with the underlying assumption that you can anthropomorphize LLMs. Gradient descent and backpropagation don't take place in the brain. LLMs "learn" in the same way that Excel sheets "learn". Humans are living beings with needs and rights. A person being able to legally squat in a home doesn't mean that a drone occupying p…

sure, but if I use an LLM to write a novel/article, I can be sued in civil court not the LLM. but, more importantly, OpenAI can also be sued for tortious interference? (basically the civil equivalent of accessory)

Whoever operates the LLM, in this case OpenAI, engaged in copyright infringement through the unauthorized modification, reproduction and distribution of content to you.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#312

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

Probably because the contents are what's posted, i.e. if someone would post a link to an interesting video behind paywall / login and there was an easy mirror available that'd be posted too.

If I could just buy one article for a coffee without entering a bunch of PII or go through a time-wasting process I would agree on the moral equivalence between the examples.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#313
post #283

Earlier quoted context omitted.

It is am ethical grey area, but if the paywall applied to all user agents, which would make it similar to say buying a Kindle book, then you might see that as pirating, whereas if you use an archive service that was served the HTTP response and cached it, then you are using a proxy UA. If the news/magazine doesn't want this they can simple serve a cut down or zero length article to all non-paying viewers! But they wa…

We can extend this analogy. What if someone put up a proxy, that has a legal Netflix subscription and which "watches" streams of Netflix shows, captures actual RGB values of pixels and re-streams the resulting video to anyone else? Isn't it the same "proxy" excuse?

I would say no because the site was happy to serve the content publicly, whereas your proxy is breaking a contractual agreement. Now we get into terms of service of a website, and even if you visit for free you agree to them. Which is a possible point. It is quite grey IMO. In terms of HN I reckon a mag would love the free brand rec. vs. the archive not being shared. Where it hurts them is if someone is avoiding paying for a subscription by continually using archive sites.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#314

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

If ChatGPT is based on neural networks, with no actual save-and-replicate facsimile behaviour, it no more "copies" original work than I do when I tell you about the news article I read today.

I'd say the only real reason the Piratebay links thing you mentioned is not the norm is purely because those media sources have done a better job of striking fear into people doing that, so it's gone more underground. I.e. they're better terrorists.

There's no fundamental, moral reason why Piratebay links being posted and raised to the top would be wrong.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#315
post #187

Earlier quoted context omitted.

In that case, the language model calls a search function and just repeats the result out its conversation context, not its training data. With that in mind it's not clear why it's ok for Bing itself to quote the source, but it stops being ok, when a chatbot does it.

Bing links to the source, chatbots doesn't.

In the example from the article, it very clearly points to all the sources used.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#316

We developers like to pretend that LLM's are akin to humans and that they've been using things like NYTimes like humans as educational material. But they are not. It's much simpler, proprietary writing is now integrated into the source code of OpenAI, it would be as if I would copy parts of other propriety code and copy paste it into my own codebase. Claiming copy paste is a natural evolving process of millions of ye…

> it would be as if I would copy parts of other propriety code and copy paste it into my own codebase. It's not copy-pasted; it's compressed in a lossy manner. Even GPT4 has nowhere near enough memory to store the entirety of its training data in a non-lossy compression format. Just likes how humans compress the information we read.

So if compress nytimes articles into a vector database and query it is a vector then that's okay in line with your reasoning?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#317

Earlier quoted context omitted.

It's inevitable that this question ends up at the supreme court. And the sooner the better IMO. It's clearly fair use. Generative agents will be seen legally as no different than a human artist leveraging the summation of their influences to create a new work.

Clearly fair use? What if I pay ChatGPT to give me the NYT article it sourced verbatim as stored (i.e. without referring me to the NYT source)?

It's not stored in ChatGPT actually, unlike Google's web search cache where it is stored verbatim, can be recalled perfectly, and is still fair use.

Fair use has nothing to do with reproducibility. LLMs are more clearly fair use than a search engine cache and those court cases are long settled. There's no world in which OpenAI doesn't win this entire thing.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#318
post #278

Earlier quoted context omitted.

Good comment, it was very funny to see how people desperately try to find moral justification for pirating media A but not B. "It's apples to oranges, you see, there are less letters in the NYT article than in the book and they are rendered differently, so it is ok to pirate their work. I did nothing wrong!" :)

There's no way to get your money back if you didn't like the content. If they don't want their articles to be read for free then they should keep them out of my view. And certainly not use clickbaity headlines. Information can be copied and they should accept it, or change their business/distribution model.

So if I went to a cinema and didn't like the movie, I should be entitled for a return, right? Or if I went into a museum and didn't like the art displayed there?

If you are advocating for a free for all libertarian dystopia, well, I have some bad news for you - they never work.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#319

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

Oh it's worse than that. The NYT is positing that any neural network that is trained on their data, and can summarize or very closely approximate an article's content on request, is in violation.

This reasoning would presumably apply to any neural network, including one made of neurons, dendrites, and axons. So any human reader of the NYT who is capable of accurately summarizing what they read is an evil copyright violator, and must be "deleted".

Effectively, the NYT legal department is setting the stage for mass murder.

Post reply on HN