Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

231–240 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#231

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

This suggests to me that copyright laws are becoming out of date. The original intent was to provide an incentive for human authors to publish work, but has become more out of touch since the internet allowed virtually free publishing and copying. I think with the dawn of LLMs, copyright law is now mainly incentivising lawyers.

> The original intent was to provide an incentive for human authors to publish work, but has become more out of touch since the internet allowed virtually free publishing and copying. I think with the dawn of LLMs, copyright law is now mainly incentivising lawyers.

And yet the content industry still creates massive profits every year from people buying content.

I think internet-native people can forget that internet piracy doesn’t immediately make copyright obsolete simply because someone can copy an article or a movie if sufficiently motivated. These businesses still exist because copyright allows them to monetize their work.

Eliminating copyright and letting anyone resell or copy anything would end production of the content many people enjoy. You can’t remove content protections and also maintain the existence of the same content we have now.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#232

This is just rent seeking from dying media instead of working on creating something new in my view. AI indeed is reading and using material sa a source, but is deriving results based on that material. I think this should be allowed, but now it is a fight who has better paid politicians pretty much. I am open to hear other thoughts.

Here’s another thought: It’s good that there are real incentives to produce original content. Especially investigative journalism which is an extremely tough business financially — even without LLMs — but with lots of social value. It would be silly to totally destroy the incentive to produce new technologies like LLMs, but so wouldn’t it be silly to destroy the incentive to produce original, high-quality content eit…

The whole idea of "a dying media" is pretty scary to me. It indicates that some people place no value in journalism. To be fair, there are a huge number of newspapers who also place little to no value in journalism. I have a number of local papers who will report on celebrity gossip, but it's all auto-translate from somewhere and just posted without questioning, so you end up with random "news" about a person who is completely unknown in the country.

Real, and especially investigative, journalism is extremely expensive and it's not something modern AI is even remotely capable to doing. It might be able to help and make it cheaper, but you can't replace newspapers with ChatGPT and expect to get anything but random gossip and rehashed press releases. I do wonder why the New York Times believe you can.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#233
post #123

Google can look up into their index and can remove whatever they want to, within minutes. But how that can be possible for an LLM? That is, "decontaminate" the model from certain parts of the corups? I can only think of excluding the data set from the training and then retrain? As a side note, I think LLM frenzy would be dead in few years, 10 years time frame at max. The rent seeking on these LLMs as of today would n…

Unix kinda still does the same thing now as before.

Future big ai models might be totally different in quality, and latency.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#234
And here I am thinking it'd be amazing to have an AI that can on-demand read me every novel ever written. It'd be even cooler to jump into a text adventure game of any novel and have it actually follow the original text.

I guess that clashes with our copyright world. (Is there hope of some kind of Netflix/Spotify model, with fractional royalties?)

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#235

Earlier quoted context omitted.

> Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. Foreign companies can be barred from selling infringing products in the United States. Russian and Chinese consumers are less interested in English-language articles. I can’t really get beh…

>Chinese LLMs are clearly trained to avoid answering certain topics that their government deems sensitive But they're not; you can download open source Chinese base models like Yi and Deepseek and ask them about Tianmen Square yourself and see, they don't have any special filtering.

I suspect they will crack down on that within the next few years.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#236

Earlier quoted context omitted.

If LLMs actually create added value and don't just burn VC money then they should be able to pay a fair price for the work of people they're relying upon. If your business is profitable only when you get your raw materials for free it's not a very good business.

By that logic you should have to pay the copyright holder of every library book you ever read, because you could later produce some content you memorised verbatim.

The rules we have now were made in the context of human brains doing the learning from copyrighted material, not machine learning models. The limitations on what most humans can memorize and reproduce verbatim are extraordinarily different from an LLM. I think it only makes sense to re-explore these topics from a legal point of view given we’ve introduced something totally new.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#237

Earlier quoted context omitted.

Here’s another thought: It’s good that there are real incentives to produce original content. Especially investigative journalism which is an extremely tough business financially — even without LLMs — but with lots of social value. It would be silly to totally destroy the incentive to produce new technologies like LLMs, but so wouldn’t it be silly to destroy the incentive to produce original, high-quality content eit…

How are the LLMs rent seeking? they are clearly providing value that people want to pay for..

This has to be one of the most abused terms on this website.

“People are willing to pay for it” is not even relevant to the question of whether it’s rent-seeking. Rent-seeking has to do with capturing unearned wealth, i.e. taking someone else’s work and profiting from it.

There is some portion of OAI’s (et al.) value that they themselves produce. There is another portion that is totally derivative of the data — other people’s work — they have trained on for free. A simple thought experiment can tell you to what degree OAI et al are “rent-seekers.”

Imagine a world where they had to enter into mutual agreements in order to train on that data. How much would the AI companies be worth? Not quite zero, but fairly close (Andreessen pretty much stated this IIRC). How much would the data producers be worth? The exact same amount or more.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#238

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

They probably didn’t start with a lawsuit. They started asking for royalties. They probably didn’t get an offer they thought was fair and reasonable so they sued. These media businesses have shareholders and employees to protect. They need to try and survive this technological shift. The internet destroyed their profitability but AI threatens to remove their value proposition.

Sorry, how exactly LLM threatens NYT? Are people supposed to generate news themselves? Or like wait a year or so before NYT articles are consumed by LMMs?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#240

Earlier quoted context omitted.

I don't think they're looking to prevent the inevitable, but rather see a target with a fat wallet from which a lot of money can be extracted. I'm not saying this in a negative way, but much of the "this is outrageous!" reaction to AI hasn't been about the building of models, but rather the realization that a few players are arguably getting very rich on those models so other people want their piece of the action.

If this is inevitable (and I'm not saying it's not), who will produce high quality news content?

AI. And, I fear, it will be good.
Post reply on HN