Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

101–110 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#101
The way to view this kind of parasitism is how we look at patent trolls. When you look at the RIAA/MPAA lawsuits, while I don't agree with them, at least file sharing was basically a canonical form of copyright infringement.

With LLMs we have an aspect of a text corpus that the creators were not using (the language patterns) and had no plans for or even idea that it could be used, and then when someone comes along and uses it, not to reproduce anything but to provide minute iterative feedback in training, they run in to try and extract some money. It's parasitism. It doesn't benefit society, it only benefits the troll, there is no reason courts should enforce it.

Someone should try and show that a NYT article can be generated autoregressively and argue it's therefore not copyrightable.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#102
post #21

For me it's quite obvious that if you make a profit from an engine that has as an input copyrighted material, then you owe something to the owner of this copyrighted content. We have seen this same problem with artists claiming stable diffusion engines were using their art.

Do all automakers that now develop electric cars owe Tesla something as they cashed in once they saw Tesla's successful copyrighted material l? A model is semantic, it contains the idea which is not copyrightable. Only how it is expressed could be copyrighted (i.e. if it outputs the copyright work verbatim). If this were not the case we would have plenty of monopolies and the world would fall apart.

If Tesla thinks their competitors have violated any of their parents they are well within their rights to seek damages...

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#103
post #71
post #12

Can someone explain the technical difference between what search engines do to index newspapers versus what is being claimed here? Is the difference as simple as me being able to get summaries and content from a newspaper from GPT without needing to visit their website?

Imagine a paid streaming service that has, say, the Lord of the Rings trilogy as part of their catalogue. They’ll be happy if you send people searching for “watch Lord of the Rings now” to their landing page. But if instead you send everyone who searches for that an .mkv of Lord of the Rings that’s ripped from their site, they’ll probably be be less happy.

Wha rid that .mkv was actually a high quality reenactment with different actors and millions of slight differences peppered throughout the story. And if a viewer of the original and a viewer of the .mkv talked about the movie they would agree on most things. But the color of the sunset or the home town name of the main character maybe different?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#104

The arguments about being able to mimic New York Times “style” are weak, but the fact that they got it to emit verbatim NY Times content seems bad for OpenAI: > As outlined in the lawsuit, the Times alleges OpenAI and Microsoft’s large language models (LLMs), which power ChatGPT and Copilot, “can generate output that recites Times content verbatim

Sarah Silverman is claiming the same thing about her book. But I've tried really hard to get ChatGPT to output sentences verbatim from her book and just can't get it to. In fact, I can't even get it to answer simple questions about facts that are in her book but nowhere else -- it just says it doesn't know. Similarly I haven't been able to reproduce any text in the NYT verbatim unless it's part of a common quote or p…

it is in the legal complaint - they have ten examples of direct content. I think they got very skilled people to work on producing the evidence.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#105

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

The war on drugs has also been unwinnable from the start and yet they built an economy on top of it, with entire agencies and a prison industry. When it comes to the fabrication and exploitation of illegality, unwinnability may be a feature, not a bug.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#106

Earlier quoted context omitted.

Sarah Silverman is claiming the same thing about her book. But I've tried really hard to get ChatGPT to output sentences verbatim from her book and just can't get it to. In fact, I can't even get it to answer simple questions about facts that are in her book but nowhere else -- it just says it doesn't know. Similarly I haven't been able to reproduce any text in the NYT verbatim unless it's part of a common quote or p…

The complaint has specific examples they got from ChatGPT. There is a precedent: There were some exploit prompts that could be used to get ChatGPT to emit random training set data. It would emit repeated words or gibberish that then spontaneously converged on to snippets of training data. OpenAI quickly worked to patch those and, presumably, invested energy into preventing it from emitting verbatim training data. It…

[deleted]

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#107

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

Any piece of pie deemed too big for one person to eat will be split accordingly.

I don’t think NYT, or any other industry, for that matter knows AI isn’t going away: in fact, they likely prefer it doesn’t, so long as they can get a slice of that pie.

That’s what the WGA and SAG struck over, and won protections ensuring AI enhanced scripts or shows will not interfere with their royalties, for example.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#108

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

This suggests to me that copyright laws are becoming out of date. The original intent was to provide an incentive for human authors to publish work, but has become more out of touch since the internet allowed virtually free publishing and copying. I think with the dawn of LLMs, copyright law is now mainly incentivising lawyers.

What incentive do people have to publish work if their work is going to primarily be consumed by a LLM and spat out without attribution at people who are using the LLM?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#109
Interesting.

I think the appropriation, privatization, and monetization of "all human output" by a single (corporate) entity is at least shameless, probably wrong, and maybe outright disgraceful.

But I think OpenAI (or another similar entity) will succeed via the Sackler defense - OpenAI has too many victims for litigation to be feasible for the courts, so the courts must preemptively decide not to bother with compensating these victims.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#110

You do copyright for content that you invented and which didn't exist before. But NYT content is reporting on events truthfully to the public without any fiction or lies. Since there can be only one truth it should not matter whether NYT or Washington Post or ChatGPT is spinning it out. Unless NYT is claiming they don't report truth and publishes fiction. That is of concern since, NYT claims to reporth news truthfull…

> You do copyright for content that you invented and which didn't exist before.

I dont think that's accurate.

The Copyright Act, § 103, allows copyright protection for "compilations (of facts)", as long as there is some "creative" or "original" act involved in developing the compilation, such as in the selection (deciding which facts to include or exclude) and arrangement (how facts are displayed and in what order).

Post reply on HN