Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

241–250 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#241

Earlier quoted context omitted.

If LLMs actually create added value and don't just burn VC money then they should be able to pay a fair price for the work of people they're relying upon. If your business is profitable only when you get your raw materials for free it's not a very good business.

What is a fair price? The entire NYT library would be a fraction of a fraction of the training set (presumably).

What if even though it's a small portion of the training data, their content has an outsized influence on the output being generated? A random NYT article about Donald Trump and a random Wikipedia article about some obscure nematode might be around the same share of training data but if 10,000x more users are asking about DJT than the nematode, what is fair? Obviously they'll need to pay royalties on the usage! /s

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#242

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

Trying to prevent AI from learning from copyrighted content would look completely stupid in a decade or two when we have AIs that are just as capable as humans, but solely due to being made of silicon rather than carbon are banned from reading any copyrighted material. Banning a synthetic brain from studying copyrighted content just because it could later recite some of that content is as stupid as banning a biologic…

We have this now with humans. I've been in a lifelong sruggle for knowledge and tools that I can afford.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#243

Earlier quoted context omitted.

It's impossible to "steal" intellectual property without some kind of mind wiping device.

You must have used that device if you're making that argument in good faith.

Okay, so how is it possible to take and deprive the author of their original? The correct term would be "unauthorised copying".

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#244

Earlier quoted context omitted.

Not sure where you're coming from in this. A NYT article, once written is copyrighted. Using the content without attribution is at best plagiarism, and spitting it out the way the LLMs do is definitely a violation of if that copyright. Unless you're telling me ChatGPT has eyes and sources just like the NYT and is worrying events as it sees them too?

I don't understand. So if New York times reported on a new laws of physics and put as an article will became copyrighted? Nobody would be able to talk about it and has to discover it by themselves? How is reporting on an event different from reporting on discovering a scientific law?

The exact words used to explain the scientific law are copyrighted by the writer (presumably the paper's authors). Rephrasings are not copywrited by the source, but by the rephrasing entity (e.g. the NYT, or a teacher that made a handout for their class).

Copyright on scientific papers is most definitely a thing, by the way.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#245

Earlier quoted context omitted.

Here’s another thought: It’s good that there are real incentives to produce original content. Especially investigative journalism which is an extremely tough business financially — even without LLMs — but with lots of social value. It would be silly to totally destroy the incentive to produce new technologies like LLMs, but so wouldn’t it be silly to destroy the incentive to produce original, high-quality content eit…

How are the LLMs rent seeking? they are clearly providing value that people want to pay for..

The LLMs don't create new content, they can only rehash existing content (in term news), which they then don't pay for... It's not really the definition of rent seeking though, they do provide some value and mostly without manipulation, not deliberate anyway. LLM also aren't harmful as such, they can be, if we use them wrong, but that's not really the fault of the technology.

I do find is a bit dishonest when they charge for their services, but don't wish to pay the people who's work the models are based on. Why should I pay to use ChatGPT, if they won't pay to use my blog posts?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#246
post #132
post #93

Earlier quoted context omitted.

What? This is about whether one country wants to cede a massive economic advantage to another country.

So the US should stop enforcing copyright or child labor laws because some other countries may not, giving them an economic advantage?

In contrast to child labor laws, which are intended and written to protect vulnerable people from exploitation, current copyright laws are tailored to the interests of Disney et al.

If they were watered down, I wouldn't see any moral or ethical loss in that.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#247
post #123

Google can look up into their index and can remove whatever they want to, within minutes. But how that can be possible for an LLM? That is, "decontaminate" the model from certain parts of the corups? I can only think of excluding the data set from the training and then retrain? As a side note, I think LLM frenzy would be dead in few years, 10 years time frame at max. The rent seeking on these LLMs as of today would n…

The argument may be that having very large models that everyone uses is a bad idea, and that companies and even individuals should instead be empowered to create their own smaller models, trained on data they trust. This will only become more feasible as the technology progresses.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#248

I'd bet they win, but how do you possibly measure the dollar amount? If you strip out 100% of NYT content from GPT-4, I don't think you'd notice a difference. But if you go domain by domain and continue stripping training data, the model will eventually get worse.

I think the only feasible outcome of the NYT winning would be a royalty structure that would have OpenAI paying the NYT to access their work, including back payments

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#249
post #101

The way to view this kind of parasitism is how we look at patent trolls. When you look at the RIAA/MPAA lawsuits, while I don't agree with them, at least file sharing was basically a canonical form of copyright infringement. With LLMs we have an aspect of a text corpus that the creators were not using (the language patterns) and had no plans for or even idea that it could be used, and then when someone comes along an…

But it’s theirs, they created it and should therefore benefit from it. I’m honestly shocked at how much these companies are getting away with. It’s piracy on a massive scale. You can get a little discombobulated reading the comments from the nerds / subject idiots on this site.

Don't forget that (almost assuredly) some percentage of HN comments are made by these very LLMs in question!!

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#250

I'd bet they win, but how do you possibly measure the dollar amount? If you strip out 100% of NYT content from GPT-4, I don't think you'd notice a difference. But if you go domain by domain and continue stripping training data, the model will eventually get worse.

If developers didn't win over github / microsoft copilot, what makes you think NYT will win?

There is something that doesn't smell right with microsoft, hopefully NYT will help expose it, wich i greatly doubt

Post reply on HN