Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

721–730 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#721

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

> To keep things simple, let's say I never regurgitate chunks of verbatim NYT articles, maybe quite short snippets. You just described Google. When you think about it, it's surprising that Google is legal. However, it is well established that what Google does is perfectly legal. Remember that internally Google keeps and uses complete verbatim copies of every web page they index. Yes, Google offers a link to the sourc…

You took that quote out of context and missed the broader point in the process. The snippets provided in regular search results cannot generally replace the substance of the full articles they link to, while that's the whole point of GP's hypothetical website—it simply doesn't reproduce large chunks of text verbatim, presumably to avoid copyright infringement claims in the hypothetical's frame, and in GP's rhetorical frame to present an analogy with the information-laundering powers of LLMs that their creators claim make their exploitation of unlicensed training data fair use.

The whole point of a search engine (as we've classically known them) is to index the web and respond to queries with a list of links that you will inspect and click through on. The whole point of an LLM chatbot tool is to eliminate those inspecting and clicking-through steps, becoming a one-stop shop for content whose substance was created by someone else. That's also the whole point of GP's hypothetical, which is why it works as an analogy.

---

There are substantially better arguments for search engines being legitimate fair use. Consider, for example, transformation. AI defenders will argue that these systems are transformative because they reshuffle elements of their input in their output, but that's clearly a much weaker form of transformation than one in which the transformed work has an entirely different nature and purpose, i.e. search engines vs. the results they return. Ultimately these technicality-based "nuh uh" arguments aren't going to save the practice of training AI on unlicensed data, because they are incompatible with the spirit of copyright law even if the novel nature of these technologies means the letter of said law can't quite nail them down yet.

If these arguments do succeed, it will be because the judicial/regulatory environment in which they were applied has been corrupted by capital.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#722
post #584

Earlier quoted context omitted.

> A sibling comment mentions search engines. I think there's a big difference. A search engine doesn't replace the source, not at all. Google has been accused for years of replacing sources with their "One Box"--the big answers at the top of the page, which are usually pulled from or corroborated by search results. They don't want you to leave the search results page (where the ads are).

Google is very careful to license all the content that shows up in that interface. They even pay Wikipedia, despite legally not needing to at all.

While paying for CC content is not 100% anathema, it'd be pretty weird if that was what they were doing.

Instead, I think they're paying for this:

https://enterprise.wikimedia.com/

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#723

Earlier quoted context omitted.

That's a good bet. Down at the bottom of the linked PDF are some more interesting allegations: Count 5 - MS/OpenAI removed NYT copyright notices in violation of the DMCA. Count 7 - By attributing hallucinated garbage to NYT, MS/OpenAI is diluting NYT trademarks in violation of US Trademark law. I admit: I laughed. This will be an entertaining lawsuit to follow.

What will ultimately happen is that OpenAI and all big tech with have to pay out some sizable sum to large copyright holders, and in exchange be granted a de facto exclusive right to develop these technologies further because they’re the only ones who can do so “responsibly” with respect to copyright. It will take a long time to wind its way through the courts, but this could be the death knell for open source LLMs i…

Meanwhile, open source LLMs are excluded from stringendo regulation in the US, abd with Mistral there is some knowhow that isn't in SV, which is also jicem

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#724

Earlier quoted context omitted.

>But if you simply ask for a summary of a specific article by, say, just name and date, and you get a copy of it, it's clear that GPT is storing the original data in some way, and thus it has copied the NYT's protected works without permission. In this particular case they were using it via Bing, which actively did a HTTP request to the particular article to extract the content. So GPT hadn't memorised it verbatim, i…

The article states that they used it initially through ChatGPT, but that seems to have been fixed in the meantime, at least for the very simplistic queries that used to work ("the first paragraph of the Carl Zimmer article on old DNA" in ChatGPT used to return the exact data from NYT, and "next paragraph" could then be used to get the following ones). Even if this has been fixed, it still proves that ChatGPT encodes…

If it had exact copies they would have showed it could recall the 8th paragraph or something. Even google and the nyt release the first paragraph for free.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#725
post #361
post #194

Earlier quoted context omitted.

What will happen in this case is that large content providers will get paid directly and smaller content providers will get rolled up into a licensing bag and get small indirect payouts. For example, we might see a model where people who's books have been used will get a pay out proportionate to the sales of the book (perhaps), so if your books sells just a few thousand copies expect $20 but if you sell millions expe…

> large content providers will get paid directly I'm sure that's what they want, but I'm not sure that's what the outcome will be. What if they want to charge a prohibitive amount of money for their content?

Dunno - but my guess is that the price will be what the market will bear...

I think Spotify vs Napster is a good example, content creators in news (the Journalists) are already in a hard place (vs. successful rock stars preinternet) I think that the news providers are rather like the music lables.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#726
post #642

Earlier quoted context omitted.

> it's clearly not helping their market value if people are checking on ChatGPT instead of reading a NYT article. People are not using ChatGPT as a replacement for current news, and because of hallucinations, no one should be using it for past news either. I wouldn't remotely call ChatGPT a competitor of NYT traffic, like I would Reuters or other news outlets.

The intended result is clearly to supplant other information sources in favor of people getting their information from ChatGPT. Why should it matter to legality that the tech isn't good enough for the goal?

> T. Why should it matter to legality that the tech isn't good enough for the goal?

Because if it is not good enough, then it is not a market substitute.

The laws cares if it is a market substitute and if there are damages. If it sucks, then there aren't damages, which matters for the 4th factor of fair use.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#727

Earlier quoted context omitted.

Copyright law doesn't mention opt outs or search engine snippet controls. It's not clear to me that robots.txt is the singular thing that makes Google legal. In US copyright law facts cannot be copyrighted, so copyright on factual content like newspaper articles is limited. Simply replacing a few words wouldn't work, but I am certain that GPT-4 is capable of paraphrasing factual content at a level that would not be c…

>Copyright law doesn't mention opt outs or search engine snippet controls. It's not clear to me that robots.txt is the singular thing that makes Google legal. Genuinely - what are you talking about besides your own assumptions? you just assume everything google does is legal and therefore any one else doing anything arguably similar must also be legal? Without regard for factual details that do matter to copyright la…

> you just assume everything google does is legal

Not an assumption. This is well established. They've been doing it for twenty years!

> Without regard for factual details that do matter to copyright law? Such as license??

What license? Google doesn't in general have or need an explicit license to crawl websites and neither does OpenAI.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#728

Earlier quoted context omitted.

I see the exact opposite - any open source model is going to become prohibitively expensive to train if quality data costs billions of dollars. We’re going to be left with the OpenAI’s and Google’s of the world as the only players in the space until someone solves synthetic data.

Exactly this. I work at a small web scraping company (so I might be a bit bias) and any small business can collect a fair, capable datasets of public data for model training, sentiment analysis or whatever today. If public data is stopped by copyright as this lawsuit implies that would just mean only giant corporations and pirates would be able to afford this. This would be a huge blow to open-source and research dev…

research is fair use, also providing something amazing like Wikipedia is arguably educational (again fair use), reselling NYT articles on-demand via an API is by itself neither, so likely not free use

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#729

Earlier quoted context omitted.

That's a good bet. Down at the bottom of the linked PDF are some more interesting allegations: Count 5 - MS/OpenAI removed NYT copyright notices in violation of the DMCA. Count 7 - By attributing hallucinated garbage to NYT, MS/OpenAI is diluting NYT trademarks in violation of US Trademark law. I admit: I laughed. This will be an entertaining lawsuit to follow.

What will ultimately happen is that OpenAI and all big tech with have to pay out some sizable sum to large copyright holders, and in exchange be granted a de facto exclusive right to develop these technologies further because they’re the only ones who can do so “responsibly” with respect to copyright. It will take a long time to wind its way through the courts, but this could be the death knell for open source LLMs i…

The prompts shown literally invite the LLM to complete the copyrighted text by providing unedited selections and asking the machine to finish those. Even if this is problematic in a small number of cases it is not a use case that undermines the business model of the newspaper since it requires the reader to have access to the original text. Nor will it be easy to demonstrate economic harm since this is not how readers consume news and is very far from how users interact with LLMs. Nor are the archival materials used for training remotely reflective of the "time-sensitive" articles that newspapers sell. And archival materials are easily available elsewhere so where is the case for economic harm?

The courts are going to rule that LLM training is a transformative use case that is protected as fair use under copyright law. They may rule that if an LLM-powered service is explicitly designed to enable copyright violation that is illegal, but there is no way any court is going to look at these examples and see it as anything other than the NYT fishing to try and generate a violation by using the LLM in a way that is very different than the service is intended to be used and which -- even if abused -- doesn't hurt the business model under which the text has been produced.

The most likely outcome is that LLM providers will add some sort of filter on output to prevent machines from regurgitating source documents. But this isn't a court case the NYT can win without gutting fair use protections, and that would be a terrible thing.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#730

Earlier quoted context omitted.

No. In the US, whether or not you make money has little to do with whether or not your use qualifies as "fair use".

Why do you say that? Commercial vs noncommercial use is a primary factor in the “purpose” prong of the fair use balancing test and a significant one in the “market effects” prong. That a use is noncommercial is often a deciding factor in the success of a fair use defense. GP is overstating it though, since it’s still one of many factors.

Your parent is more right than you.

Weird Al has made a fantastic living copying music while only changing lyrics. He makes very heavy use of the satire plank of Fair Use.

The “commercial” test is only part of the decision criteria for Fair Use.

Post reply on HN