Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

681–690 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#681

> Because the outputs of Defendants’ GenAI models compete with and closely mimic the inputs used to train them, copying Times works for that purpose is not fair use. This is interesting. The NYT is specifically saying that the way you use an LLM impacts what you can legally use for training the LLM. They're firing shots at the big guys trying to sell access to an LLM, but not at the little guy self-hosting for fun or…

They're probably saying that because its what the supreme court said except about a human copying a work created by another human.

https://www.npr.org/2023/05/18/1176881182/supreme-court-side...

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#682
post #670

Earlier quoted context omitted.

it's fair use if you don't make money from your project no?

No. In the US, whether or not you make money has little to do with whether or not your use qualifies as "fair use".

Why do you say that? Commercial vs noncommercial use is a primary factor in the “purpose” prong of the fair use balancing test and a significant one in the “market effects” prong.

That a use is noncommercial is often a deciding factor in the success of a fair use defense. GP is overstating it though, since it’s still one of many factors.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#683

Earlier quoted context omitted.

If they could find a single person who in natural use (e.g. not as they were trying to gather data for this lawsuit) has ever actually used ChatGPT as a direct substitution for a NYT subscription, I'd support this lawsuit. But nobody would do that, because ChatGPT is a really shitty way to read NYT articles (it's stale, it can't reliably reproduce them, etc.). All that is valuable about it is the way that it transfor…

That’s nonsense piracy. I never intend to own a truck, so when I need to haul a little something I go to Home Depot and steal a Ford off the lot for an hour? What if I stole all your commits, plucked the hard lines out of the ceremony, and then launched an equivalent feature the same week as you did, but for a competing software company? Would you or your employer deserve to get paid for my use of the slice of your w…

awful comparison

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#684
post #677

Earlier quoted context omitted.

> I don't think the lawsuit has any merit The lawsuit fundamentally has merit. It asks a huge open question that no one knows the answer to. The outcome will be extraordinarily impactful. The question must be answered at some point. The case has merit even if NYT loses across the board.

How is the question being asked different from the Google Books case?

I don't think anybody has claimed that OpenAI is causing NYT subscriptions to go up. NYT has even expressly made the claim they're losing potential revenue.

> [1] On the most important factor, possible economic damage to the copyright owner, [Judge] Chin wrote that "Google Books enhances the sales of books to the benefit of copyright holders."

[1]: https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#685

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

What you described is entirely fair use, actually. Not only that, look at a few news articles from Tier 2 and down publications, and you'll realize that almost all of them are directly sourced from NYT and others. They'll say "so and so happened, according to The Times" (and usually link the article there)

>What you described is entirely fair use, actually.

Based upon what? You think other publishers use NYTimes articles for free without license?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#686

Earlier quoted context omitted.

No. In the US, whether or not you make money has little to do with whether or not your use qualifies as "fair use".

Why do you say that? Commercial vs noncommercial use is a primary factor in the “purpose” prong of the fair use balancing test and a significant one in the “market effects” prong. That a use is noncommercial is often a deciding factor in the success of a fair use defense. GP is overstating it though, since it’s still one of many factors.

Because anyone that is familiar with fair use knows that the purpose prong and the commerciality aspect of it is not one of the more important prongs of the fair use analysis, whereas transformation is. Transformation adjusts what is a purpose that falls under fair use. Did you read Warhol??

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#687

I think there is a national security aspect to ML models trained on copyrighted data. Countries that allow it will gain a superior technological advantage and outcompete those who disallow training on copyrighted material. I personally believe training LLMs on copyrighted data is copyright infringement if the models are deployed in a way that competes with the copyright holder. But that doesn’t necessarily mean it’s…

You can say the same for any legal enforcement like respecting patent or copyright law or making Champagne outside France. Yet the sky isn’t falling given this reality with so many legally protected industries. Maybe these markets where such an industry might offshore to are too small and insular to be very significant, and are probably language bound to make english models less relevant compared to native language m…

Champagne isn’t a transformative technology, and least not anymore.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#688
post #625

Earlier quoted context omitted.

Any publisher can opt out of google. Publisher also have substantial control over titles and snippets shown in google, whether an article appears in google news, etc Paraphrasing is also known as cloning and is often a copyright violation

Copyright law doesn't mention opt outs or search engine snippet controls. It's not clear to me that robots.txt is the singular thing that makes Google legal. In US copyright law facts cannot be copyrighted, so copyright on factual content like newspaper articles is limited. Simply replacing a few words wouldn't work, but I am certain that GPT-4 is capable of paraphrasing factual content at a level that would not be c…

>Copyright law doesn't mention opt outs or search engine snippet controls. It's not clear to me that robots.txt is the singular thing that makes Google legal.

Genuinely - what are you talking about besides your own assumptions? you just assume everything google does is legal and therefore any one else doing anything arguably similar must also be legal? Without regard for factual details that do matter to copyright law? Such as license?? Your own description of copyright law here is very stunted - you can't paraphrase articles of the NYTimes and call it a fair use. You can report on what the NYtimes reports on... because that's what news is.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#689

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

What you described is entirely fair use, actually. Not only that, look at a few news articles from Tier 2 and down publications, and you'll realize that almost all of them are directly sourced from NYT and others. They'll say "so and so happened, according to The Times" (and usually link the article there)

Do you have some examples & are you sure they don't pay licensing fees to NYT?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#690

> Because the outputs of Defendants’ GenAI models compete with and closely mimic the inputs used to train them, copying Times works for that purpose is not fair use. This is interesting. The NYT is specifically saying that the way you use an LLM impacts what you can legally use for training the LLM. They're firing shots at the big guys trying to sell access to an LLM, but not at the little guy self-hosting for fun or…

They're probably saying that because its what the supreme court said except about a human copying a work created by another human. https://www.npr.org/2023/05/18/1176881182/supreme-court-side...

That's a good bet.

Down at the bottom of the linked PDF are some more interesting allegations:

Count 5 - MS/OpenAI removed NYT copyright notices in violation of the DMCA.

Count 7 - By attributing hallucinated garbage to NYT, MS/OpenAI is diluting NYT trademarks in violation of US Trademark law.

I admit: I laughed. This will be an entertaining lawsuit to follow.

Post reply on HN