Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

341–350 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#341
post #327

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

Why can't AI at least cite its source? This feels like a broader problem, nothing specific to the NYTimes. Long term, if no one is given credit for their research, either the creators will start to wall off their content or not create at all. Both options would be sad. A humane attribution comment from the AI could go a long way - "I think I read something about this in the NYTimes on January 3rd, 2021." It appears t…

A neural net is not a database where the original source is sitting somewhere in an obvious place with a reference. A neural net is a black box of functions that have been automatically fit to the training data. There is no way to know what sources have been memorized vs which have made their mark by affecting other types of functions in the neural net.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#342
post #123

Google can look up into their index and can remove whatever they want to, within minutes. But how that can be possible for an LLM? That is, "decontaminate" the model from certain parts of the corups? I can only think of excluding the data set from the training and then retrain? As a side note, I think LLM frenzy would be dead in few years, 10 years time frame at max. The rent seeking on these LLMs as of today would n…

> But how that can be possible for an LLM?

They should have thought of that before they went ahead and trained on whatever they could get.

Image models are going to have similar problems, even if they win on copyright there's still CSAM in there: https://www.theregister.com/2023/12/20/csam_laion_dataset/

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#343

Earlier quoted context omitted.

Doesn't this harm open source ML by adding yet another costly barrier to training models?

It doesn't matter what's good for open source ML. It matters what is legal and what makes sense.

Clearly if a law is bad then we should change that law. The law is supposed to serve humanity and when it fails to do so it needs to change.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#345

Earlier quoted context omitted.

Doesn't this harm open source ML by adding yet another costly barrier to training models?

It doesn't matter what's good for open source ML. It matters what is legal and what makes sense.

It matters what ends up being best for humanity, and I think there are cases to be made both ways on this

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#346

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

The way I see it, if the NYT goes under (one of the biggest newspapers in the world), all similar outlets also go under. Major publishers, both of fiction and non-fiction, as well as images, video, and all other creative content, may also go under. Hence, there is no more (reliable) training data.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#347

Earlier quoted context omitted.

Doesn't this harm open source ML by adding yet another costly barrier to training models?

It doesn't matter what's good for open source ML. It matters what is legal and what makes sense.

Slavery was legal...

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#348
post #168

Earlier quoted context omitted.

I disagree. A printer is too neutral - it's just a tool, like roads or the internet. Third parties can use them to commit copyright infringement, but that doesn't (or shouldn't) reflect on the seller of the tool. I propose it's more like selling a music player that comes preloaded with (remixes of) recording artists' songs.

It is neutral though. That’s the whole point. You have to twist its arm with great intention to recreate specific things. Sufficient intention that it’s really on you at that point.

It’s not neutral if all the content is in the model, regardless of whether you had to twist its arm or not. What does that even mean with a piece of software?

A printer is neutral because you have to send it all the data to print out a copy of copyrighted content. It doesn’t contain it inherently.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#349
post #170

Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.) I don’t necessarily fault OpenAI’s decision to initially train their models without entering into licensing agreements - they probably wouldn’t exist and the generative AI revolution may never have happened if…

It’s likely fair use.

What if a court interprets fair use as a human-only right, just like it did for copyright?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#350
post #170

Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.) I don’t necessarily fault OpenAI’s decision to initially train their models without entering into licensing agreements - they probably wouldn’t exist and the generative AI revolution may never have happened if…

Doesn't this harm open source ML by adding yet another costly barrier to training models?

I think not, because stealing large amounts of unlicensed content and hoping momentum/bluster/secrecy protects you is a privilege afforded only to corporations.

OSS seems to be developing its own, transparent, datasets.

Post reply on HN