Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

31–40 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#31

Does anyone know what the copyright status of LLM generated content is? That is, if I feed a NYT article into GPT4 and say, summarize this article, and then publish that summary, is there argument or precedent that says that is or is not copyright infringement? Asking for a friend.

[deleted]

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#32
post #29

The challenge for all these AI companies is that the only thing of value for building a defensible commercial product is having proprietary datasets for training. With the underlying techniques and algorithms all being rapidly commoditized the power lies in who holds and owns that data. Like all other ML “revolutions” it’s the training data that matters and if one doesn’t have access to training data others don’t hav…

And I imagine that Gmail makes google very very special in this regard

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#35
post #22
post #3

> The New York Times is suing OpenAI and Microsoft over claims the companies built their AI models by “copying and using millions” of the publication’s articles and now “directly compete” with the outlet’s content. Millions? Damn, they can churn out some content. 13 million[0]!. [0] https://archive.nytimes.com/www.nytimes.com/ref/membercenter... .

The NYT publishes about 200 pieces of journalism every day (according to their own website), and it was founded in 1851. That makes for a lot of articles.

(2023 - 1851) * 365 * 200 = 12,556,000

Yep, so a few million ripped off articles is plausible.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#36
I think the train has left the station and the ship has sailed. I'm not sure it's possible to put this genie back in the bottle. I had stuff stolen by OpenAI too, and I felt bad about it (and even send them a nasty legal letter when it could output my creative work almost verbatim), but I think at this point, the legal landscape needs to somehow adjust. The Copyright Clause in the US Constitution is clear:

To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries

Blocking LLMs on the basis of copyright infringement does NOT promote progress in science and the useful arts. I don't think copyright is a useful basis to block LLMs.

They do need to be regulated, and quickly, but that regulatory regime should be something different. Not copyright. The concept of OpenAI before it became a frankenmonster for-profit was good. Private failed, and we now need public.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#37

The arguments about being able to mimic New York Times “style” are weak, but the fact that they got it to emit verbatim NY Times content seems bad for OpenAI: > As outlined in the lawsuit, the Times alleges OpenAI and Microsoft’s large language models (LLMs), which power ChatGPT and Copilot, “can generate output that recites Times content verbatim

I'm not sure if the verbatim content isn't more of a "stopped clock is right twice a day" or "monkeys typewriting shakespeare" situation. As I see it, most of the value in something like the NYT is as a trusted and curated source of information with at least some vetting. The content regurgitated from an LLM would be intermixed with false information and all sorts of other things, none of which are actually news from a trusted source - the main reason people subscribe to the NYT (?) and something at which ChatGPT cannot directly compete against NYT writers.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#38
post #8

What are they arguing here? AFAIK reading copyrighted works is not copyright infringement. Copying and selling them is, as the name would suggest, but OpenAI absolutely did not do that. Are they trying to say that LLM training is a special type of reading that should be considered infringement? Seems like a weak case to me. edit: Would be very funny if OpenAI used an educational fair use defense

If you read the complaint, you will see that, among others, paragraphs and paragraphs of NYT articles are reproduced verbatim or almost verbatim.

Sounds like the infinite monkey typewriter thing, where the NYT sieves the OpenAI output for exact segments, probably after primining it to fatten the yield

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#39

The arguments about being able to mimic New York Times “style” are weak, but the fact that they got it to emit verbatim NY Times content seems bad for OpenAI: > As outlined in the lawsuit, the Times alleges OpenAI and Microsoft’s large language models (LLMs), which power ChatGPT and Copilot, “can generate output that recites Times content verbatim

I would agree. Style is too amorphous (even among its own reporters and journalists, there are different styles), but verbatim repetition would be a problem. So what would the licensing be for all their content be (if presumably one could get ChatGPT to output all of the NYTs articles)?

The unfortunate thing about these LLMs is they siphon all public data regardless of license. I agree with data owners one can’t Willy nilly use data that’s accessible but not licensed properly.

Obviously Wikipedia, data from most public institutions, etc., should be available, but not data that does not offer unrestricted use.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#40

What are they arguing here? AFAIK reading copyrighted works is not copyright infringement. Copying and selling them is, as the name would suggest, but OpenAI absolutely did not do that. Are they trying to say that LLM training is a special type of reading that should be considered infringement? Seems like a weak case to me. edit: Would be very funny if OpenAI used an educational fair use defense

The second paragraph of the article is > As outlined in the lawsuit, the Times alleges OpenAI and Microsoft’s large language models (LLMs), which power ChatGPT and Copilot, “can generate output that recites Times content verbatim, closely summarizes it, and mimics its expressive style.” This “undermine[s] and damage[s]” the Times’ relationship with readers, the outlet alleges, while also depriving it of “subscription…

> closely summarizes it

Absolutely not copyright infringement

> mimics its expressive style

Absolutely not copyright infringement

> can generate output that recites Times content verbatim

This one seems the closest to infringement, but still doesn't seem like infringement. A printer has this capability too. If a user told ChatGPT to recite NYT content and then sold that content, that would be 100% infringement, but would probably be on the user, not the tool. e.g. if someone printed out NYT articles and sold them, nobody would come after the printer manufacturer.

> undermine[s] and damage[s]” the Times’ relationship with readers, the outlet alleges, while also depriving it of “subscription, licensing, advertising, and affiliate revenue.

This claim seems far fetched as the point of the NYT is to report the news. One thing that LLMs absolutely cannot do is report today's news. I can see no way that ChatGPT is a substitute for the NYT in a way that violates copyright.

Post reply on HN