Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

191–200 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#191

Earlier quoted context omitted.

“Through Microsoft’s Bing Chat (recently rebranded as “Copilot”) and OpenAI’s ChatGPT, Defendants seek to free-ride on The Times’s massive investment in its journalism by using it to build substitutive products without permission or payment,” the lawsuit states. I can't be the only one that sees the irony of this news being "reported" and regurgitated over dozens of crappy blogs. ChatGPT [..] “can generate output tha…

> More on-topic: if the NYT thinks that GPT-4 is replicating their style then [as anybody who has tried to do creative writing work can testify to] they need to fire all their writers. The complaint isn’t that ChatGPT is imitating New York Times style by default. The complaint is that you can ask it to write “in the style of New York Times” and it will do so. I don’t know if this argument has any legal merit, but it’…

The writing of the new york times is so diffuse (they even have their own published style guide!) that it's impossible to make a claim to any "style", as there are undoubtedly millions upon millions of lines of text by authors who have been inspired by the NYT.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#192

Earlier quoted context omitted.

If NYT wins this, then there is going to be a massive push for payouts from basically everyone ever…I don’t see that wallet being fat for long.

If LLMs actually create added value and don't just burn VC money then they should be able to pay a fair price for the work of people they're relying upon. If your business is profitable only when you get your raw materials for free it's not a very good business.

By that logic you should have to pay the copyright holder of every library book you ever read, because you could later produce some content you memorised verbatim.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#193
For many of our google searches, the first results tend to be wikipedia, instagram, etc... We click on those clicks and both google and the clicked website get a share of our traffic. So it is somewhat fair.

But in current AI situation, wikipedia, nytimes, stackoverflow etc are getting a pretty unfair deal. Probably all major text based outlets are seeing a drop in their numbers now...

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#194

Earlier quoted context omitted.

This suggests to me that copyright laws are becoming out of date. The original intent was to provide an incentive for human authors to publish work, but has become more out of touch since the internet allowed virtually free publishing and copying. I think with the dawn of LLMs, copyright law is now mainly incentivising lawyers.

What incentive do people have to publish work if their work is going to primarily be consumed by a LLM and spat out without attribution at people who are using the LLM?

I think you're making a profound point here.

I believe you equate incentive to monetary rewards. And while that it probably true for the majority of news outlets, money isn't always necessarily what motivates journalists.

So considering the hypothetical situation where journalists (or more generally, people that might publish stuff) were somehow compensated. But in this hypothetical, they would not be attributed (or only to very limited extent) because LLMs are just bad at attribution.

Shouldn't in that case the fact that information distribution by the LLM were "better" be enough to satisfy the deeper goal of wanting to publish stuff? Ie.: reach as many people looking for that information as possible, without blasting it out or targeting and tracking audiences?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#195
> “copying and using millions” of the publication’s articles and now “directly compete” with its content as a result.

The New York Times doesn't have a lot faith in the quality of their own content. How on earth is ChatGPT going to go out into the world a do reporting from Gaza or Ukraine? How is it going to go to the presidents press conference and ask questions? ChatGPT cannot produce original content in the same way a newspaper can. The fact that the NYT seems to believe that ChatGPT can compete says a lot about how they write their articles or their lack of understanding of how LLMs work.

Now I do believe that OpenAI could at least have asked the newspapers before just scraping their content, but I think they knew that that would have undermined their business model, which tells you something about how tech companies work.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#196

What next, suing the school system for using NYT articles in English class to train children?

If the school is selling access to those articles and/or passing the information off as their own original content? Yes.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#197
post #36

I think the train has left the station and the ship has sailed. I'm not sure it's possible to put this genie back in the bottle. I had stuff stolen by OpenAI too, and I felt bad about it (and even send them a nasty legal letter when it could output my creative work almost verbatim), but I think at this point, the legal landscape needs to somehow adjust. The Copyright Clause in the US Constitution is clear: To promote…

I see, the narrative switched form “cat’s out of the bag” to “genie’s out of the bottle”. Regardless, no one wants to ban llms. We just want the theft to stop.

it's not theft

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#198
post #3

> The New York Times is suing OpenAI and Microsoft over claims the companies built their AI models by “copying and using millions” of the publication’s articles and now “directly compete” with the outlet’s content. Millions? Damn, they can churn out some content. 13 million[0]!. [0] https://archive.nytimes.com/www.nytimes.com/ref/membercenter... .

“Through Microsoft’s Bing Chat (recently rebranded as “Copilot”) and OpenAI’s ChatGPT, Defendants seek to free-ride on The Times’s massive investment in its journalism by using it to build substitutive products without permission or payment,” the lawsuit states. I can't be the only one that sees the irony of this news being "reported" and regurgitated over dozens of crappy blogs. ChatGPT [..] “can generate output tha…

I'm pretty sure the defense is the "NYT Style and Usage Guide"...

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#199

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

I don't think they're looking to prevent the inevitable, but rather see a target with a fat wallet from which a lot of money can be extracted. I'm not saying this in a negative way, but much of the "this is outrageous!" reaction to AI hasn't been about the building of models, but rather the realization that a few players are arguably getting very rich on those models so other people want their piece of the action.

If this is inevitable (and I'm not saying it's not), who will produce high quality news content?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#200

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

> Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. Foreign companies can be barred from selling infringing products in the United States. Russian and Chinese consumers are less interested in English-language articles. I can’t really get beh…

>Chinese LLMs are clearly trained to avoid answering certain topics that their government deems sensitive

But they're not; you can download open source Chinese base models like Yi and Deepseek and ask them about Tianmen Square yourself and see, they don't have any special filtering.

Post reply on HN