Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

1–10 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#3
> The New York Times is suing OpenAI and Microsoft over claims the companies built their AI models by “copying and using millions” of the publication’s articles and now “directly compete” with the outlet’s content.

Millions? Damn, they can churn out some content. 13 million[0]!.

[0] https://archive.nytimes.com/www.nytimes.com/ref/membercenter....

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#4
What are they arguing here? AFAIK reading copyrighted works is not copyright infringement. Copying and selling them is, as the name would suggest, but OpenAI absolutely did not do that. Are they trying to say that LLM training is a special type of reading that should be considered infringement? Seems like a weak case to me.

edit: Would be very funny if OpenAI used an educational fair use defense

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#5
I'd bet they win, but how do you possibly measure the dollar amount? If you strip out 100% of NYT content from GPT-4, I don't think you'd notice a difference. But if you go domain by domain and continue stripping training data, the model will eventually get worse.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#6

What are they arguing here? AFAIK reading copyrighted works is not copyright infringement. Copying and selling them is, as the name would suggest, but OpenAI absolutely did not do that. Are they trying to say that LLM training is a special type of reading that should be considered infringement? Seems like a weak case to me. edit: Would be very funny if OpenAI used an educational fair use defense

>AFAIK reading copyrighted works

I hope you don’t think that’s all whats happening, right?

>LLM training is a special type of reading that should be considered infringement

OK, what turn of phrase would you prefer?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#7

What are they arguing here? AFAIK reading copyrighted works is not copyright infringement. Copying and selling them is, as the name would suggest, but OpenAI absolutely did not do that. Are they trying to say that LLM training is a special type of reading that should be considered infringement? Seems like a weak case to me. edit: Would be very funny if OpenAI used an educational fair use defense

The article mentions that ChatGPT will absolutely parrot back NYTimes article text verbatim. So yes, it's copyright infringement.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#8

What are they arguing here? AFAIK reading copyrighted works is not copyright infringement. Copying and selling them is, as the name would suggest, but OpenAI absolutely did not do that. Are they trying to say that LLM training is a special type of reading that should be considered infringement? Seems like a weak case to me. edit: Would be very funny if OpenAI used an educational fair use defense

If you read the complaint, you will see that, among others, paragraphs and paragraphs of NYT articles are reproduced verbatim or almost verbatim.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#9
post #2

NYT article with a lot more context https://www.nytimes.com/2023/12/27/business/media/new-york-t...

But what's the ChatGPT summary, in an expressive style that closely mimics a reader's personal relationship to the Times?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#10

What are they arguing here? AFAIK reading copyrighted works is not copyright infringement. Copying and selling them is, as the name would suggest, but OpenAI absolutely did not do that. Are they trying to say that LLM training is a special type of reading that should be considered infringement? Seems like a weak case to me. edit: Would be very funny if OpenAI used an educational fair use defense

The second paragraph of the article is

> As outlined in the lawsuit, the Times alleges OpenAI and Microsoft’s large language models (LLMs), which power ChatGPT and Copilot, “can generate output that recites Times content verbatim, closely summarizes it, and mimics its expressive style.” This “undermine[s] and damage[s]” the Times’ relationship with readers, the outlet alleges, while also depriving it of “subscription, licensing, advertising, and affiliate revenue.”

Post reply on HN