Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

741–750 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#741

Earlier quoted context omitted.

They're probably saying that because its what the supreme court said except about a human copying a work created by another human. https://www.npr.org/2023/05/18/1176881182/supreme-court-side...

That's a good bet. Down at the bottom of the linked PDF are some more interesting allegations: Count 5 - MS/OpenAI removed NYT copyright notices in violation of the DMCA. Count 7 - By attributing hallucinated garbage to NYT, MS/OpenAI is diluting NYT trademarks in violation of US Trademark law. I admit: I laughed. This will be an entertaining lawsuit to follow.

Very interested how this turns out as IIUC copyright violations have statutory damages which the NYT won't have to prove.

$750 [1] * 66 million records [the lawsuit] is basically 50 billion.

[1]: https://www.ce9.uscourts.gov/jury-instructions/node/706

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#742
post #670

Earlier quoted context omitted.

it's fair use if you don't make money from your project no?

No. If that were true, I could take a band that I hate, copy all of their music note-for-note, then release an exact copy on the market and undercut them by selling their entire discography for $0.01 Fair Use requires one of several enumerated activities, including satire, education, journalism. You can’t just copy content and hope that it passes Fair Use. Hire a lawyer if you are unsure. But at least read the Wikipe…

Cover songs do get a compulsory license though, for a predetermined royalty, and one condition is not changing it too much.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#743
post #446
post #419

Earlier quoted context omitted.

I'm sorry but this is such a bad take. Nice appeal to consequences. In my view, the New York Times is entirely justified in pursuing legal action. They invested time and effort in creating content, only to have it used without permission for monetary gain. A clear violation. Analyzing the factors involved for a "fair use" consideration: Purpose and Character of the Use: While the argument for transformation might hol…

I don’t think the original point being made was that NYT wasn’t justified in bringing the action. The point that was being made was the suit would be ultimately meaningless in the long term even if it was successful in the short term. There is a potentially more significant risk in the future that this suit will not protect against because of the reasons enumerated by the author. While the author is speculating, the…

Correct, a_wild_dandan argues that the outcome of this suit makes no pragmatic difference.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#744
post #674

Earlier quoted context omitted.

No, in US law at least there can be no copyright of facts, only presentation. If you convey the same facts in different words that isn't a matter of fair use, it's never even a matter of copyright in the first place.

How about things that aren’t quite facts? Reviews, opinions, etc.

Illegal to have the same opinion as someone?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#745

Earlier quoted context omitted.

I agree, but nothing worth having is free. NYT and other news outlets have to ultimately pay reporters to go out into the world and do the work. The reporters are not priests, and the NYT is not a church that lives off donations and tax exemptions. They need money to operate, and you may disagree with how they try to collect that money (paywall) but that doesn't solve their funding problem. How would you pay for news…

> How would you pay for news otherwise? You could subsidise news via "public service" style stipends. Much like having a government owned "independent" news service (eg the BBC) this comes with a high risk of corruption. Don't bite the hand that feeds and all that. You could implement a much lower friction non-recurring payment system. I'd be far more tempted to drop a little money on a fixed term (5 articles, 1 day,…

> I'd be far more tempted to drop a little money on a fixed term (5 articles, 1 day, ???) setup than a subscription.

I believe that’s often referred to as a newspaper, which should be available in all good newsagents on any given day.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#746

The lawsuit itself (which arstechnica links to): https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20... From page 30 and onwards has some fairly clear examples on how ChatGPT has an (internal) copy of copyrighted material which it will recite verbatim. Essentially if you copy a lot of copyrighted material into a blob and then apply some sort of destructive compression to it. How destructive would that compre…

I wonder how they got these results, seeing as they are not showing any of the usual UI's (i.e. ChatGPT/Copilot).

It makes it difficult for me to ascertain whether it is repeating from it's training data, or they committed the same mistake as the OP article of using Copilot, which ends up googling(binging?) the article first, before replying.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#747

Earlier quoted context omitted.

> OpenAI could declare that a sufficient degree of AGI has been achieved and thus return to its philanthropic mission The response from MSFT's legal team would be biblical if openai pulled this.

It’s literally in the contract such a distinction is at OpenAIs discretion

The lawsuit won't be about the clause, it would be about the definition of AGI.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#748

Earlier quoted context omitted.

> What you described is entirely fair use, actually Just like during the pandemic how everyone became an epidemiologist, suddenly everyone's a copyright lawyer. I'll just dispute your assertion by saying: 1. Questions of fair use are famously gray, and anyone who declares something as "entirely fair use", with no caveats, is nearly always wrong except for the must obvious cases, which the given example is most defini…

> suddenly everyone's a copyright lawyer Roll back 20+ years ago on Slashdot and you'll see the exact same thing. Copyright has been a hot button issue on the internet for decades. People end up thinking (rightly or wrongly) that they understand it without being a lawyer.

> that they understand it without being a lawyer.

Quite literally, not even the lawyers or courts understand it. This is very much a "learn as you go" exercise for humanity in general at this point in time.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#749

Earlier quoted context omitted.

What you described is entirely fair use, actually. Not only that, look at a few news articles from Tier 2 and down publications, and you'll realize that almost all of them are directly sourced from NYT and others. They'll say "so and so happened, according to The Times" (and usually link the article there)

> What you described is entirely fair use, actually Just like during the pandemic how everyone became an epidemiologist, suddenly everyone's a copyright lawyer. I'll just dispute your assertion by saying: 1. Questions of fair use are famously gray, and anyone who declares something as "entirely fair use", with no caveats, is nearly always wrong except for the must obvious cases, which the given example is most defini…

> Questions of fair use are famously gray, and anyone who declares something as "entirely fair use", with no caveats, is nearly always wrong except for the must obvious cases, which the given example is most definitely not. A judge has wide latitude in determining fair use.

You're the one presenting unfounded claims with confidence here. There is well established case law about not being able to copyright facts. If you are actually fully paraphrasing a presentation of facts / ideas and not just altering a couple of words here and there, then there is a very strong case for non-infringement.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#750

Earlier quoted context omitted.

> To keep things simple, let's say I never regurgitate chunks of verbatim NYT articles, maybe quite short snippets. You just described Google. When you think about it, it's surprising that Google is legal. However, it is well established that what Google does is perfectly legal. Remember that internally Google keeps and uses complete verbatim copies of every web page they index. Yes, Google offers a link to the sourc…

You took that quote out of context and missed the broader point in the process. The snippets provided in regular search results cannot generally replace the substance of the full articles they link to, while that's the whole point of GP's hypothetical website—it simply doesn't reproduce large chunks of text verbatim, presumably to avoid copyright infringement claims in the hypothetical's frame, and in GP's rhetorical…

A search engine takes an input string, a corpus of text, and returns a series of text that best comes next after the input string.

An LLM takes an input string a corpus of text, and returns a series of text that best comes after the next input string.

To get a paragraph of output, you run the search over and over again

Both the search and LLM reshuffle the inputs to the outputs.

If I'm describing the purpose of the LLM, it's got a wide number of usages. "Making my resume look more professional" or "be a crud api" or "reformat my ask into a api call to X service" or "give me a timeline of events surrounding Y with source links"

Post reply on HN