Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

41–50 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#41
post #25
post #12

Can someone explain the technical difference between what search engines do to index newspapers versus what is being claimed here? Is the difference as simple as me being able to get summaries and content from a newspaper from GPT without needing to visit their website?

NYTimes is an ad-supported business, so you visiting their website to read the content those ads pay for is important.

NYT is an anomaly. The majority of their revenue is actually from subscriptions.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#42
post #17

Earlier quoted context omitted.

Take estimated losses of the NYT from this "innovation" and multiply by 10^x where is "x" high enough to make tech companies stop and think before they break laws next time. That would be my approach at least.

which laws are broken exactly? it's not remotely settled law that "training an NN = copyright infringement"

The training isn't the issue per se, it's the regurgitation of verbatim text (or close enough to be immediately identifiable) within a for-profit product. Worse still that the regurgitation is done without attribution.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#43
post #3

> The New York Times is suing OpenAI and Microsoft over claims the companies built their AI models by “copying and using millions” of the publication’s articles and now “directly compete” with the outlet’s content. Millions? Damn, they can churn out some content. 13 million[0]!. [0] https://archive.nytimes.com/www.nytimes.com/ref/membercenter... .

“Through Microsoft’s Bing Chat (recently rebranded as “Copilot”) and OpenAI’s ChatGPT, Defendants seek to free-ride on The Times’s massive investment in its journalism by using it to build substitutive products without permission or payment,” the lawsuit states. I can't be the only one that sees the irony of this news being "reported" and regurgitated over dozens of crappy blogs. ChatGPT [..] “can generate output tha…

> More on-topic: if the NYT thinks that GPT-4 is replicating their style then [as anybody who has tried to do creative writing work can testify to] they need to fire all their writers.

The complaint isn’t that ChatGPT is imitating New York Times style by default.

The complaint is that you can ask it to write “in the style of New York Times” and it will do so.

I don’t know if this argument has any legal merit, but it’s not as simple as you suggest. It’s the textual parallel to having AI image generators mimic the trademark style of artists. We know it can be done, the question is what does it mean legally.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#44
post #16

What are they arguing here? AFAIK reading copyrighted works is not copyright infringement. Copying and selling them is, as the name would suggest, but OpenAI absolutely did not do that. Are they trying to say that LLM training is a special type of reading that should be considered infringement? Seems like a weak case to me. edit: Would be very funny if OpenAI used an educational fair use defense

> AFAIK reading copyrighted works is not copyright infringement. [...] Are they trying to say that LLM training is a special type of reading that should be considered infringement? Nobody can argue that OpenAI was feeding the content to ChatGPT because ChatGPT was bored or was curious about current events. It was fed NYT's content so it would know how to reproduce similar content, for profit. I think getting a case-l…

> ... for profit.

for non-profit

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#45
post #17

Earlier quoted context omitted.

Take estimated losses of the NYT from this "innovation" and multiply by 10^x where is "x" high enough to make tech companies stop and think before they break laws next time. That would be my approach at least.

which laws are broken exactly? it's not remotely settled law that "training an NN = copyright infringement"

The legal argument, which I'm sure you are very well aware of, is that training a model on data, reorganizing, and then presenting that data as your own is copyright infringement.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#46
post #17

Earlier quoted context omitted.

Take estimated losses of the NYT from this "innovation" and multiply by 10^x where is "x" high enough to make tech companies stop and think before they break laws next time. That would be my approach at least.

which laws are broken exactly? it's not remotely settled law that "training an NN = copyright infringement"

Right, hence the lawsuit. They allege that the Copyright Act is the law that was broken.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#47
post #22

Earlier quoted context omitted.

The NYT publishes about 200 pieces of journalism every day (according to their own website), and it was founded in 1851. That makes for a lot of articles.

(2023 - 1851) * 365 * 200 = 12,556,000 Yep, so a few million ripped off articles is plausible.

Everything from 1851 to 1927 ought to be in the public domain, though. If the goal of training an AI is just "to mimic a style" there are absolutely humongous amounts of text that's totally free of any copyright restrictions.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#48
post #6

What are they arguing here? AFAIK reading copyrighted works is not copyright infringement. Copying and selling them is, as the name would suggest, but OpenAI absolutely did not do that. Are they trying to say that LLM training is a special type of reading that should be considered infringement? Seems like a weak case to me. edit: Would be very funny if OpenAI used an educational fair use defense

>AFAIK reading copyrighted works I hope you don’t think that’s all whats happening, right? >LLM training is a special type of reading that should be considered infringement OK, what turn of phrase would you prefer?

Of course, OpenAI doesn't just read. But do they simply reproduce content verbatim? And when they do not reproduce contents of the NYT verbatim, but rather process and tailor them to the situation, mixed, 'charged' with new 'information content' and adjusted to the purpose of the inquirer, it will not be easy for the New York Times.

Because ultimately, our entire knowledge is based on the knowledge of others and is remixed, 'charged' and changed by us after reading. I also think that the New York Times uses the contents of others to create new content.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#49

Does anyone know what the copyright status of LLM generated content is? That is, if I feed a NYT article into GPT4 and say, summarize this article, and then publish that summary, is there argument or precedent that says that is or is not copyright infringement? Asking for a friend.

No one knows, this is new territory. Maybe the fermi filter is litigating an AI that would otherwise save humanity.

> an AI that would otherwise save humanity.

Just to clarify, this is sarcasm right?

Post reply on HN