Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

151–160 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#151

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

All of this can be true (I don’t think it necessarily is, but for the sake of argument), but it’s legally irrelevant: the court is not going to decide copyright infringement cases based on geopolitical doctrines.

Courts don’t decide cases based on whether infringement can occur again, they decide them based on the individual facts of the case. Or equivalently: the fact that someone will be murdered in the future does not imply that your local DA should not try their current murder cases.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#152

Earlier quoted context omitted.

This suggests to me that copyright laws are becoming out of date. The original intent was to provide an incentive for human authors to publish work, but has become more out of touch since the internet allowed virtually free publishing and copying. I think with the dawn of LLMs, copyright law is now mainly incentivising lawyers.

What incentive do people have to publish work if their work is going to primarily be consumed by a LLM and spat out without attribution at people who are using the LLM?

I would guess the monetisation is going to be limited to either subscriptions or advertising if your reputation allows people to especially value your curation of facts/reporting etc. The big issue with LLMs is the lack of reliability - it might be accurate or it might be an hallucination.

Personally, I think it would be a lot simpler if the internet was declared a non-copyright zone for sites that aren't paywalled as there's already a legal grey area as viewing a site invariably involves copying it.

Maybe we'll end up with publishers introducing traps/paper towns like mapmakers are prone to do. That way, if an LLM reproduces the false "fact", it'll be obvious where they got it from.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#153

Earlier quoted context omitted.

This suggests to me that copyright laws are becoming out of date. The original intent was to provide an incentive for human authors to publish work, but has become more out of touch since the internet allowed virtually free publishing and copying. I think with the dawn of LLMs, copyright law is now mainly incentivising lawyers.

What incentive do people have to publish work if their work is going to primarily be consumed by a LLM and spat out without attribution at people who are using the LLM?

To have a positive impact on the world? Also, presumably NYT still has a business model unrelated to whatever OpenAI is doing with their data and everyone working there is still getting paid for their work...

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#154
post #123

Google can look up into their index and can remove whatever they want to, within minutes. But how that can be possible for an LLM? That is, "decontaminate" the model from certain parts of the corups? I can only think of excluding the data set from the training and then retrain? As a side note, I think LLM frenzy would be dead in few years, 10 years time frame at max. The rent seeking on these LLMs as of today would n…

But, wasn’t the reason proprietary unixes died out at major work horses because of a nearly feature comparable free alternative (Linux)?

Extending the analogy, LLMs won’t die out, just proprietary ones. (Which is where I think this tech will actually go anyway.)

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#155

Earlier quoted context omitted.

This suggests to me that copyright laws are becoming out of date. The original intent was to provide an incentive for human authors to publish work, but has become more out of touch since the internet allowed virtually free publishing and copying. I think with the dawn of LLMs, copyright law is now mainly incentivising lawyers.

What incentive do people have to publish work if their work is going to primarily be consumed by a LLM and spat out without attribution at people who are using the LLM?

Without 200 years of copyright protection, how will any author be able to afford food?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#156

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

I see a complete economic collapse unless creators start getting paid both for their data upfront, and paid royalties when their data is used in an LLM response

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#157
post #109

Interesting. I think the appropriation, privatization, and monetization of "all human output" by a single (corporate) entity is at least shameless, probably wrong, and maybe outright disgraceful. But I think OpenAI (or another similar entity) will succeed via the Sackler defense - OpenAI has too many victims for litigation to be feasible for the courts, so the courts must preemptively decide not to bother with compen…

What does the Sackler defence refer to?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#158
post #123

Google can look up into their index and can remove whatever they want to, within minutes. But how that can be possible for an LLM? That is, "decontaminate" the model from certain parts of the corups? I can only think of excluding the data set from the training and then retrain? As a side note, I think LLM frenzy would be dead in few years, 10 years time frame at max. The rent seeking on these LLMs as of today would n…

You may be interested in https://unlearning-challenge.github.io/

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#159
post #25
post #12

Can someone explain the technical difference between what search engines do to index newspapers versus what is being claimed here? Is the difference as simple as me being able to get summaries and content from a newspaper from GPT without needing to visit their website?

NYTimes is an ad-supported business, so you visiting their website to read the content those ads pay for is important.

My browser doesn’t display ads and shows little regard for most paywalls. Do I owe NYTimes something?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#160
post #99
post #21

For me it's quite obvious that if you make a profit from an engine that has as an input copyrighted material, then you owe something to the owner of this copyrighted content. We have seen this same problem with artists claiming stable diffusion engines were using their art.

If you study copyrighted material for four years at a university and then go on to earn money based on your education, do you owe something to the authors of your text books? I'm not sure how we should treat LLMs with respect to publicly accessible but copyrighted material, but it seems clear to me that "profiting" from copyrighted material isn't a sufficient criteria to cause me to "owe something to the owner".

We don’t, and shouldn’t, give LLMs the same rights as people.
Post reply on HN