Live data from Hacker News

New York Times considers legal action against OpenAI as copyright tensions swirl

npr.org

131–140 of 383 posts

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#131
post #90

> if a federal judge finds that OpenAI illegally copied The Times' articles to train its AI model, the court could order the company to destroy ChatGPT's dataset, forcing the company to recreate it using only work that it is authorized to use. I'd like to see it happening but it sounds unrealistic.

They would simply move AI training operations to other jurisdictions where it's legal.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#132

Earlier quoted context omitted.

> we end up with an inferior product cannibalizing a superior one and driving it out of business. In case that print is meant by inferior product: The same argument could've been brought up for Napster, where traditional distribution via CD printing through music labels are the inferior product driving the superior one out of business. Or rather it's big labels suing Napster out of business. I also hold a dislike for…

For me the question is highly debatable. As far as I'm aware the training of AI works by crawling various content from the Internet, so they're using a product, which are NY articles to train an AI, meaning they're using their content to help creating a product. But in this sense, shouldn't they be suing Google as well? Since Google as a search engine, also crawls the web and shows their articles in their search resu…

> But in this sense, shouldn't they be suing Google as well? Since Google as a search engine, also crawls the web and shows their articles in their search results, usually it may even use them for those quick answers features.

Not sure if NYT is involved as a plaintiff, but this has happened in Europe: https://en.wikipedia.org/wiki/Ancillary_copyright_for_press_...

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#133

Don't humans operate similarly? We gain knowledge through experiences. These AI models effectively condense a vast amount of experience data into weights. Considering the global race in AI advancements, I'm skeptical about the success of these copyright claims. I do find it hypocritical that OpenAI says that other LLMs can't be trained on data generated by their LLMs.

Not just humans in general, the NYT specifically relies on the fact news cannot be copyrighted. Then claims it's articles are sacred...

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#134

Earlier quoted context omitted.

But OpenAI is neither copying nor deriving a work. Style (which can be described as a probability model) is not copyrightable. You're halfway there with #2. The output is not copyrightable, but unless you can actually point to a sequence of words from the original it can't be infringing.

> But OpenAI is neither copying nor deriving a work. They absolutely are copying in the course of training the model, and they are doing something that often looks a lot like copying when producing output with the model. > Style (which can be described as a probability model) is not copyrightable. Style is not all that can be described in an LLM's "probability model", otherwise LLM models would never be able to repro…

> They absolutely are copying in the course of training the model

Only in the sense that a Cisco router is copying in the course of sending me the article, which we've all agreed doesn't count as infringement.

The bigger problem is that the plaintiff has to show it is more likely than not that that sequence of words came from their text and not some other source, which is going to be obscenely difficult.

> And ChatGPT is quite capable of producing verbatim text from its training set.

Granted, and then the copyright holder could sue for infringement at that point. Exactly whom he should sue is a more difficult question.

Look, I get it: you want copyright to allow authors to say "you can't do that with my work"; I'm sympathetic, but it just doesn't give authors that power.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#135
post #36
post #9

Earlier quoted context omitted.

Paraphrasing is not the issue. The issue is that OpenAI copied the Times ’ creative works into a GPU to train a model. That copy was likely neither licensed nor fair use.

Do you think OpenAI doesn’t subscribe? I too load creative works into all sorts of temporary structures in order to read the paper. I don’t need to license it, I pay for a subscription. Should I pay more if I memorize the paper? Should I pay more if I read it to my sick friend in the hospital? Should I pay more if I save copies to my own hard drive and grep for words in the files? Should robots have a higher subscrip…

> This whole “they copied it into a gpu” doesn’t matter. People read and interpret. Robots read and interpret. I don’t want to live in a world where every specific device and use needs to be licensed. Especially ex post facto. That will suck so hard.

I think you already live in that world (though IANAL), there was a ruling that the Glider cheat tool for WoW was a copyright violation even though it was poking around inside the local copy necessarily made in RAM as part of normal usage of WoW.

https://arstechnica.com/gaming/2009/01/judges-ruling-that-wo...

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#136

Earlier quoted context omitted.

News organizations in other jurisdictions already have achieved settlements with Google (which has much deeper pockets than OpenAI) But there's a fairly obvious difference in use between using content to index it and point to it and generate revenue for it and using content to generate alternative content...

If you're using their content to generate more content, doesn't it fall under fair use?

In the case of Google, merely indexing content is not considered fair use. It's a double edged sword, as media outlets have realized, after all Google is responsible for a larger part of the success of these publications.

But in Germany for example, Google News basicially just copypasted articles into their service, and monetizing it without involving the publishers. That doesn't qualify as transformative even under their own rules (see the YouTube TOS and copyright enforcement system).

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#137

These mega LLMs that can autonomously roam the web and consume original content are basically the "I made this" meme[0] and having some legal precedent would be good for all users of the web. [0] - https://knowyourmeme.com/memes/i-made-this

> having some legal precedent would be good for all users of the web

Honestly that's true whichever way it falls. The sooner it's clear what's allowed and what's not, the better for everyone.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#139
I think that the proper outcome for all of this would be acknowledgement that the current copyright laws very poorly regulate this aspect, that the key parts of any such legal action are at the not-really-described edges of law because these edges weren't relevant until now; and so instead of waiting for courts ruling on how law-as-written-now applies and accepting these rulings, we will likely get some new legislation explicitly setting what the legal norms should be.

In the short term, of course, the existing law matters, but the main discussion should be not on how to apply existing law but how to ensure that the new laws match what we-the-people would want.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#140
post #90

> if a federal judge finds that OpenAI illegally copied The Times' articles to train its AI model, the court could order the company to destroy ChatGPT's dataset, forcing the company to recreate it using only work that it is authorized to use. I'd like to see it happening but it sounds unrealistic.

If I read 1000s of of NYT articles to improve my writing skills, add then write an article of my own, is that a copyright violation?

The problem with the "a lossy mathematical translation of its inputs is exactly like a person learns" arguments, even if courts don't find them ludicrous, is that people absolutely can and are found guilty of trademark violations when they read thousands of pages of the LOTR and then write a fantasy novel full of Tolkein's character names for profit.
Post reply on HN