Live data from Hacker News

New York Times considers legal action against OpenAI as copyright tensions swirl

npr.org

111–120 of 383 posts

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#111
post #90

> if a federal judge finds that OpenAI illegally copied The Times' articles to train its AI model, the court could order the company to destroy ChatGPT's dataset, forcing the company to recreate it using only work that it is authorized to use. I'd like to see it happening but it sounds unrealistic.

If I read 1000s of of NYT articles to improve my writing skills, add then write an article of my own, is that a copyright violation?

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#112
post #70

There is a very real risk that we end up with an inferior product cannibalizing a superior one and driving it out of business. Moreover, AI would seem to be even more susceptible to capture and manipulation than conventional media. When it's a question of guiding thought I prefer the humanities to tech. (Same with art.)

Cannibalizing is a good description, the inferior product basically only exists thanks to the superior one (remove training data and it's nothing) but also threatens to eliminate it...

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#113
Honestly, I think generative AI losing a massive copyright showdown is inevitable at this stage.

It's extremely easy to get the latest generation of AIs to produce outputs that in many fields sans-AI would be trivially considered as IP infringement.

While there are many interesting reasonable legal & technical arguments that it's not, the result completely undermines copyright protections regardless. If that's accepted at scale, copyright in practice will change completely. In effect, the choices are "block this, or entirely destroy copyright protections in many industries". You can't allow this without eventually allowing everybody to simulate their own NY Times reporters, produce their own Marvel movies, and create their own Taylor Swift albums.

If you do allow that, the many many affected industries have catastrophic problems.

Problematic though copyright laws are, I see no world where all those protections go away any time soon, and so if the courts don't agree to protect copyright already in this scenario, then it will eventually be legislated to make that happen. AI consuming copyrighted data and producing an output has to be considered a derivative work (or indeed, the model itself will be considered a derivative work) or IP protections are effectively broken.

There's a grace period now while we work our way there, but the politics is pretty clear and with no plausible path to "let's drop copyright completely" ASAP, I just don't see any other result in the medium term. Doesn't mean the end of generative AI by any means, just a slowdown as we move to a world where you need to negotiate rights and buy data to feed it first, instead of scraping everybody else's for free.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#114
post #44

Earlier quoted context omitted.

If I buy a license to listen to a Taylor Swift mp3, I can copy it onto a flash drive (or anywhere I like to use it). And that’s fine until I distribute copies to others.

Yes, and if OpenAI had a license to use the Times ’ works as model training material this would be a non-issue, of course.

I didn’t buy a license for Taylor Swift that explicitly allows me to copy to a flash drive. Or to listen with only one ear, etc.

I bought a license to listen and use.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#115

Earlier quoted context omitted.

Did the Times grant a license to every router on the internet to transmit its intellectual property to other routers? If not, the judge should grant an injunction contingent on requiring the Times to verify that every person who accesses their content is doing so only over routers and other devices with express written authorization, for every step in the process. Maybe even extend it to browsers and client libraries…

You don’t need a fair use exemption for transient copies in service of licensed or fair uses. Computers and networks have been around a long time. These issues have been given a good workout.

> You don’t need a fair use exemption for transient copies in service of licensed or fair uses.

The transient copy isn’t for a licensed use, and whether it is a Fair Use is specifically the subject of debate. So this really is basically an admission that but for the potential applicability of Fair Use, the use of the copyrighted mmaterial in training is a violation of copyright.

Also, even if the training is fair use, that doesn’t mean that the copies of the source material produced by OpenAI and distributed to their customers using the model are fair use. Just becaue making the tool is a transformative fair use doesn’t mean using the tool to generate copies of the material which was used to train it, which are significantly less trandormative than the model itself, are Fair Use. (And the fact that one of the functions that the model is used for is this commercial, for profit by the maker of the model, copying of the source material is – as much as I believe AI model training on its own is quite likely to generally be fair use – an argument against the model training being fair use in this csase.)

> Computers and networks have been around a long time.

True, and commercially producing and delivery copies of copyrighted works in a manner which substitutes for the original work in the marketplace, no matter what intermediate steps go into doing that, and no matter that computers or networks are used in those intermediate steps, is pretty much the clearest case of violation of copyright you can get.

> These issues have been given a good workout.

Some of them have, some of them have not. Whether and in what conditions training an AI on source material that may be subject in aggregate to a compilation copyright by someone else, and which consists further of individual works that have their own copyrights, might be “fair use” is not one of the issues that have been given a good workout. Neither – because producing predictive models in that way has not previously been common – has whether, unlike other intermediate tool use, using such a model in the course of doing what would otherwise be an infringement by producing a copy of specific copyright-protected works, commercially, for a customer at their request, is no longer a violation because the use of the model somehow isolates it from liability,

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#116

Earlier quoted context omitted.

Really? Seems like it's not so different conceptually to search engines. They also make money by indexing "training data" of a sort. But the open web trundles on regardless, because the search engines found ways to cut the content producers in on it. I see no reason why that can't be the case here too, with AI companies training their models to act more like search engines when data comes from certain sources - i.e.…

>Really? Seems like it's not so different conceptually to search engines. They also make money by indexing "training data" of a sort. OK, so if a writer X has a blog to put up samples of their work to drive people to buy books and to get writing assignments and someone uses ChatGPT to write something in the style of X - this naively seems like a hit on that author's ability to sell their skills. And I don't think it…

> this naively seems like a hit on that author's ability to sell their skills

If the AI is good, then the author's economic outlook is bad regardless of style.

Regardless of if the AI is or isn't good, then I can still see it being brand damaging, but my gut feeling (IANAL) is that this is more of a trademark issue than a copyright issue, as it's passing off as something it isn't.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#117
post #90

> if a federal judge finds that OpenAI illegally copied The Times' articles to train its AI model, the court could order the company to destroy ChatGPT's dataset, forcing the company to recreate it using only work that it is authorized to use. I'd like to see it happening but it sounds unrealistic.

I think it's the most natural way it would happen. Pit powerful interests (publishers) against powerful interests (Microsoft/ClosedAI). The only surprising thing is it's taking publishers so long to notice this fresh abuse of copyright at grand scale will cost them.

Which is why OpenAIs GTM strategy involves the biggest players in each industry.

That said, the outcome is unlikely - we have trained AI for more than a decade as ‘fair use’ at this point, it’s the application of the technology that is shifting the perspective, nor the act of training.

Every computer vision system in the world is trained on mostly public data for example.

Furthermore, the LLMs purpose is not to generate news so NYT will have to argue about the value of archive data. Many jurisdictions have thresholds of how much of an original work contributes to the derivative before it would be considered not fair use or plagiarism. Given the size of the datasets - good luck.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#118

Earlier quoted context omitted.

Really? Seems like it's not so different conceptually to search engines. They also make money by indexing "training data" of a sort. But the open web trundles on regardless, because the search engines found ways to cut the content producers in on it. I see no reason why that can't be the case here too, with AI companies training their models to act more like search engines when data comes from certain sources - i.e.…

>Really? Seems like it's not so different conceptually to search engines. They also make money by indexing "training data" of a sort. OK, so if a writer X has a blog to put up samples of their work to drive people to buy books and to get writing assignments and someone uses ChatGPT to write something in the style of X - this naively seems like a hit on that author's ability to sell their skills. And I don't think it…

It would work both ways. I could also imagine someone asking: “what books should I read to learn about X?” And the LLM could drive sales toward that author’s books.

It’s not clear to me that having their content “unindexed” is good for authors. It’s probably good and bad at the same time.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#119
post #90

> if a federal judge finds that OpenAI illegally copied The Times' articles to train its AI model, the court could order the company to destroy ChatGPT's dataset, forcing the company to recreate it using only work that it is authorized to use. I'd like to see it happening but it sounds unrealistic.

[deleted]

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#120
post #36

Earlier quoted context omitted.

Do you think OpenAI doesn’t subscribe? I too load creative works into all sorts of temporary structures in order to read the paper. I don’t need to license it, I pay for a subscription. Should I pay more if I memorize the paper? Should I pay more if I read it to my sick friend in the hospital? Should I pay more if I save copies to my own hard drive and grep for words in the files? Should robots have a higher subscrip…

How about “they copied it into a terrific text-to-speech engine and sold ads against the resulting podcasts”?

Reading a text has been ruled as performance and a copyright violation.

I think the issue is that LLMs don’t make a copy or distribute a copy. They use the content to create something else. I don’t remember the copyright term for whether is is transformative enough. But it basically says I can’t copy Starry Night, but I can create a painting with the same color scheme and themes as long as it’s different enough from the original.

Post reply on HN