Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

111–120 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#111
If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use?

Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NYT articles, maybe quite short snippets.

Is that fair use? IANAL, but doesn't sound like it. Typically I can't take a personal "tier" of a product and charge 3rd parties for derivatives of it. Say like VS Code.

A sibling comment mentions search engines. I think there's a big difference. A search engine doesn't replace the source, not at all. Rather it points me at it, and offers me the opportunity to pay for the article. Whereas either this or an LLM uses NYT content as an alternative to actually paying for an NYT subscription.

But then what do I know...

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#112

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

> Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee.

>Is that fair use? IANAL, but doesn't sound like it.

If you pay someone to do the summarisation for you, then you publish the content and charge a fee for it, you're the one liable, not the person you paid to summarise it for you. Similarly if you ask GPT to do it for you, then publish it, you're liable for what you publish; GPT is just a summarisation tool.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#113

Earlier quoted context omitted.

> Why someone work with full time writing articles should give the work for free OpenSource developers did that ;)

When open source developers do that, they also include an explicit licensing information that lists cases when the usage is allowed and restricted. So even if the code is open source and licensed under GPL, its usage in a closed source product like ChatGPT is not allowed.

GPL code usage in closed source ChatGPT is allowed "for internal use"; it just would not be allowed to distribute binaries of ChatGPT that are closed source without making source available; also a GPL3 license violation to allow online access to a ChatGPT program that used GPL3 code without making source available.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#114

Earlier quoted context omitted.

Many instances of fair use involve verbatim copying. The important questions surround the situation in which that happens - not so much the copying. NYT is in uncharted territory here.

in the same way that machines are not able to claim copyright, they aren't allowed to claim other legal rights either, like "fair use". The entity which owns ChatGPT is apparently maintaining a copy of the entirety of the New York Times archive within the ChatGPT knowledge base. That they extract some fair use snippets (they would claim) from it would still be fruit of a poisoned tree, no? (disclaimer: I'm pro AI, an…

There is another fix, but it will have to wait for GPT-5. They could reword articles, summarize in different words and analyze their contents, creating sufficiently different variants. The ideas would be kept, but original expression stripped. Then train GPT5 on this data. The model can't possibly regurgitate copyrighted content if they never saw it during training.

This can be further coupled with search - use GPT to look at multiple sources at once, and report. It's what humans do as well, we read the same news in different sources to get a more balanced take. Maybe they have contradictions, maybe they have inaccuracies, biases. We could keep that analysis for training models. This would also improve the training set.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#115
NYT's perspective is going to look so stupid in future when we put LLMs into mechanical bodies with the ability to interact with the physical world, and to learn/update their weights live. It would make it completely illegal for such a robot to read/watch/listen to any copyrighted material; no watching TV, no reading library books, no browsing the internet, because in doing so it could memorise some copyrighted content.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#116

Won't hold in court. GPT is a platform mainly providing answer to private individuals asking. Is like you ask a professor a question and he answered verbatim what copyrighted materials available (due to photographic memory) word for word back to you. Now if you take this answer and write a book or publish enmass on blogs for example, then you are the one should be sued by NYT. If GPT use the exact same wordings and p…

[deleted]

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#117
post #83

Earlier quoted context omitted.

I hope people start calling out the "well it's fine if a human does it" arguments out for the rat fuck thinking it is. These are computational systems operating at very large scales run by some of the wealthiest companies in the world. If I go fishing, the regulations I have to comply with are very light because the effect I have on the environment is minimal. The regulations for an industrial fishing barge are right…

unfortunately that's not the crowd of people here. 80% of the comments under this thread (right now, 2:52est) are making similar arguments and *continue* to act like LLMs are doing something unique/creative... instead of just generating sentences, from algorithms, from virtually pirated content in the form of data mining

“It is difficult to get a man to understand something, when his salary depends on his not understanding it.”

https://www.goodreads.com/quotes/21810-it-is-difficult-to-ge...

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#118
post #83

Earlier quoted context omitted.

I hope people start calling out the "well it's fine if a human does it" arguments out for the rat fuck thinking it is. These are computational systems operating at very large scales run by some of the wealthiest companies in the world. If I go fishing, the regulations I have to comply with are very light because the effect I have on the environment is minimal. The regulations for an industrial fishing barge are right…

unfortunately that's not the crowd of people here. 80% of the comments under this thread (right now, 2:52est) are making similar arguments and *continue* to act like LLMs are doing something unique/creative... instead of just generating sentences, from algorithms, from virtually pirated content in the form of data mining

“As if LLMs are doing something creative and aren’t just algorithms”

You have no idea what you’re talking about huh?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#119

NYT's perspective is going to look so stupid in future when we put LLMs into mechanical bodies with the ability to interact with the physical world, and to learn/update their weights live. It would make it completely illegal for such a robot to read/watch/listen to any copyrighted material; no watching TV, no reading library books, no browsing the internet, because in doing so it could memorise some copyrighted conte…

Will it? If the LLM in the body is allowed to read nytimes on a tablet I'm sure they wouldn't care.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#120

Companies that have content all see dollar signs. NYT won't mind if you use their content to train LLMs - as long as they get a commission. Reddit will shut down their free API and make you pay to get training content. Discord is going to be selling content for AI training too - if they haven't already done so. Twitter is doing it. They didn't care before because LLMs were just experiments. Now we're talking trillion…

NYT do not "have" content, they create content. It's their raison d'etre.
Post reply on HN