Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

21–30 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#21
For me it's quite obvious that if you make a profit from an engine that has as an input copyrighted material, then you owe something to the owner of this copyrighted content. We have seen this same problem with artists claiming stable diffusion engines were using their art.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#22
post #3

> The New York Times is suing OpenAI and Microsoft over claims the companies built their AI models by “copying and using millions” of the publication’s articles and now “directly compete” with the outlet’s content. Millions? Damn, they can churn out some content. 13 million[0]!. [0] https://archive.nytimes.com/www.nytimes.com/ref/membercenter... .

The NYT publishes about 200 pieces of journalism every day (according to their own website), and it was founded in 1851. That makes for a lot of articles.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#23

Does anyone know what the copyright status of LLM generated content is? That is, if I feed a NYT article into GPT4 and say, summarize this article, and then publish that summary, is there argument or precedent that says that is or is not copyright infringement? Asking for a friend.

No one knows, this is new territory.

Maybe the fermi filter is litigating an AI that would otherwise save humanity.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#24
post #6

What are they arguing here? AFAIK reading copyrighted works is not copyright infringement. Copying and selling them is, as the name would suggest, but OpenAI absolutely did not do that. Are they trying to say that LLM training is a special type of reading that should be considered infringement? Seems like a weak case to me. edit: Would be very funny if OpenAI used an educational fair use defense

>AFAIK reading copyrighted works I hope you don’t think that’s all whats happening, right? >LLM training is a special type of reading that should be considered infringement OK, what turn of phrase would you prefer?

You could definitely argue that it's more than just reading since they made the model out of it. But the matrix of parameters generated by training is so fundamentally different than the input that it is certainly covered by the transformative use exception to copyright.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#25
post #12

Can someone explain the technical difference between what search engines do to index newspapers versus what is being claimed here? Is the difference as simple as me being able to get summaries and content from a newspaper from GPT without needing to visit their website?

NYTimes is an ad-supported business, so you visiting their website to read the content those ads pay for is important.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#26

What are they arguing here? AFAIK reading copyrighted works is not copyright infringement. Copying and selling them is, as the name would suggest, but OpenAI absolutely did not do that. Are they trying to say that LLM training is a special type of reading that should be considered infringement? Seems like a weak case to me. edit: Would be very funny if OpenAI used an educational fair use defense

The article mentions that ChatGPT will absolutely parrot back NYTimes article text verbatim. So yes, it's copyright infringement.

Sections of this statement absolutely parrot back NYTimes article text vebatim depending how you look at it. What's the line? 3 sequential verbatim words? 5? 8?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#27
post #17

I'd bet they win, but how do you possibly measure the dollar amount? If you strip out 100% of NYT content from GPT-4, I don't think you'd notice a difference. But if you go domain by domain and continue stripping training data, the model will eventually get worse.

Take estimated losses of the NYT from this "innovation" and multiply by 10^x where is "x" high enough to make tech companies stop and think before they break laws next time. That would be my approach at least.

which laws are broken exactly? it's not remotely settled law that "training an NN = copyright infringement"

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#28

Does anyone know what the copyright status of LLM generated content is? That is, if I feed a NYT article into GPT4 and say, summarize this article, and then publish that summary, is there argument or precedent that says that is or is not copyright infringement? Asking for a friend.

Typically if you ask a chatbot to "summarize" something, it will paraphrase the original closely enough that it would be considered plagiarism and copyright infringement. To avoid that, it's required to distill the relevant ideas contained in the text, and expound on them in a way that's not dependent on how the text itself was expressed, structured or organized. You would need to tell the model to do this over multiple steps, and then derive a rephrased article without looking at the original at all. (Which is not really possible if the article was in the AI's training set, as is the case here.)

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#29
The challenge for all these AI companies is that the only thing of value for building a defensible commercial product is having proprietary datasets for training. With the underlying techniques and algorithms all being rapidly commoditized the power lies in who holds and owns that data. Like all other ML “revolutions” it’s the training data that matters and if one doesn’t have access to training data others don’t have then you’ll soon be toast.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#30
post #3

> The New York Times is suing OpenAI and Microsoft over claims the companies built their AI models by “copying and using millions” of the publication’s articles and now “directly compete” with the outlet’s content. Millions? Damn, they can churn out some content. 13 million[0]!. [0] https://archive.nytimes.com/www.nytimes.com/ref/membercenter... .

  “Through Microsoft’s Bing Chat (recently rebranded as “Copilot”) and OpenAI’s ChatGPT, Defendants seek to free-ride on The Times’s massive investment in its journalism by using it to build substitutive products without permission or payment,” the lawsuit states.
I can't be the only one that sees the irony of this news being "reported" and regurgitated over dozens of crappy blogs.

  ChatGPT [..] “can generate output that recites Times content verbatim, closely summarizes it, and mimics its expressive style.”
If the NYT thinks that GPT-4 is replicating their style then [as anybody who has tried to do creative writing work with GPT-4 can testify to] they need to fire all their writers.
Post reply on HN