Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

91–100 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#91
The lawsuit itself (which arstechnica links to):

https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20...

From page 30 and onwards has some fairly clear examples on how ChatGPT has an (internal) copy of copyrighted material which it will recite verbatim.

Essentially if you copy a lot of copyrighted material into a blob and then apply some sort of destructive compression to it. How destructive would that compression have to be for the copyright no longer to hold? My guess it would have to be a lot.

As I see it the closeness of OpenAI may be what saves it. OpenAI could filter and block copyrighted material from the LLM from leaving the web interface using some straight forward matching mechanism against the copyrighted part of the data set ChatGPT has been trained on. Whereas open source projects trained on the same data set would be left with the much harder task of removing the copyrighted material from the LLM itself.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#92

The suit demonstrates instances where ChatGTP / Bing Copilot copy from the NYT verbatim. I think it is hard to argue that such copying constitutes "fair use". However, OAI/MS should be able to fix this within the current paradigm: Just learn to recognize and punish plagiarism via RLHF. However, the suit goes far beyond claiming that such copying violates their copyright: "Unauthorized copying of Times Works without p…

Many instances of fair use involve verbatim copying. The important questions surround the situation in which that happens - not so much the copying. NYT is in uncharted territory here.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#94
post #70
post #66

Earlier quoted context omitted.

Well yeah, copying a work and using it for its original expressive purpose isn’t fair use, no? You have to use it for a transformative purpose. Suppose I’m selling subscriptions to the New Jersey Times, a site which simply downloads New York Times articles and passes them through an autoencoder with some random noise. It serves the exact same purpose as the New York Times website, except I make the money. Is that fai…

> Well yeah, copying a work and using it for its original expressive purpose isn’t fair use, no? You have to use it for a transformative purpose. They transformed the weights. Just like reading the article transforms yours . As for verbatim reproduction, I'm pretty sure brains are capable of reproducing song lyrics, musical melodies, common symbols ("cool S"), and lots of other things verbatim too. Those quotes from…

This comment is just blatant anthropomorphizing of ML models. You have no idea if reading an article “transforms weights” in a human mind, and regardless, they aren’t legally the same thing anyway.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#95
post #76
post #62

Earlier quoted context omitted.

I think the scale only matters here (probably). Because I will find it hard that a teacher/professor will not be allowed to setup a service where they will teach and provide their knowledge for others. That is basically the concept of teaching. Of course until LLM, we never had this scale before. Millions of potential learners vs the normal hundreds in a classroom session. So that makes the new case interesting

"Teaching" by copying source books word for word, would be copyright infringement; see, for example, the well-known issues around photocopying books or even excerpts. Also lying on source materials (e.g. telling students that some respected historian denies the Holocaust happened, when it's obviously not the case) is not "teaching" - it's defamation, and the NYT is absolutely right to pursue that angle too. Using LLM…

> Teaching" by copying source books word for word, would be copyright infringement; see, for example, the well-known issues around photocopying books or even excerpts.

Incorrect. Educational use helps satisfy one of tests for fair use. Teachers can, in many cases, photocopy copyrighted work without infringing on that copyright.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#96
post #84
post #72

Would be funny if NT Times won this and all commercial LLMs were shut down. Then LLMs would be distributed only via torrents, like most copyright infringing media.

Making these things anathema to commercial interests and making training them at scale legally perilous would be a huge win.

> making training them at scale legally perilous would be a huge win.

Why?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#97

I read a NYT article and publish a summary of facts that I learned: totally legit. Train a model on NYT text that outputs a summary of facts that it learned: OMG literally murder.

Fair use is intended for humans, much like copyright in general. If you can't copyright AI-generated pieces, then why would fair use apply to LLMs?

> Fair use is intended for humans.

Is it? Can you quote relevant legislation or case law?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#98
This wave is growing. Just cannot see how the big LLM players are going to get round this without paying big licence fees to content creators. Feels a bit like the torrent to Spotify moment, but for _all_ content, not just music. How they will manage the licensing model is beyond me, it’s going to be very easy for someone to sue these companies, but very difficult for the companies to calculate, attribute value and payout individual creators that contributed a tiny fraction of the training data. Surely this will make it very difficult for them to keep a business model working to a level their VC backers need to warrant even a fraction of their valuations.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#100

The suit demonstrates instances where ChatGTP / Bing Copilot copy from the NYT verbatim. I think it is hard to argue that such copying constitutes "fair use". However, OAI/MS should be able to fix this within the current paradigm: Just learn to recognize and punish plagiarism via RLHF. However, the suit goes far beyond claiming that such copying violates their copyright: "Unauthorized copying of Times Works without p…

Many instances of fair use involve verbatim copying. The important questions surround the situation in which that happens - not so much the copying. NYT is in uncharted territory here.

in the same way that machines are not able to claim copyright, they aren't allowed to claim other legal rights either, like "fair use".

The entity which owns ChatGPT is apparently maintaining a copy of the entirety of the New York Times archive within the ChatGPT knowledge base. That they extract some fair use snippets (they would claim) from it would still be fruit of a poisoned tree, no?

(disclaimer: I'm pro AI, anti copyright, especially anti elitist NY Times; but pro rule of law)

Post reply on HN