Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

661–670 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#661
post #645

I don't think the lawsuit has any merit, but I'd still like to encourage Sam Altman et al, if they really care about the greater good, to go Keyser Söze and immediately release torrents of the weights and source code for GPT-4 under GPL.

AFAIK the IP deal with Microsoft only covers development before AGI.

So at any point OpenAI could declare that a sufficient degree of AGI has been achieved and thus return to its philanthropic mission. With GPLed models and all.

However, at this point the employees expect a multi-million cash-out for each of them. So the philanthropic mission seems to be gone out the window.

And probably that’s also the way Sam Altman got back into the CEO role. By maximizing the expected eventual cash-out for the employees which threatened to leave otherwise.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#662

Earlier quoted context omitted.

Nope, it doesn't work that way. The fact that the LLM can regurgitate original articles doesn't remove the possibility that training can be considered transformative work, or more in general that using copyrighted material for training can be considered fair use. Rather, verbatim reproduction is the proof that copyrighted materials was used. Then the court has to evaluate whether it was fair use. Without verbatim rep…

What if the LLM is running locally and doing all of these things rather than hosted on a webserver which is serving the content?

It doesn't matter, if everything else stays the same what matters is what it's used for. If it's used to make money, it would certainly hurt claims of fair use—maybe not for those that do the training, but for those that use it.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#663
post #418

People who think the examples the lawsuit are “fair use” need to consider what that would mean. We’re basically going to let a few companies consolidate all the value on the Internet into their black boxes with basically no rules … that seems very dangerous to me. I hope a court establishes some rules of engagement here, even if it’s not this case.

I see the exact opposite - any open source model is going to become prohibitively expensive to train if quality data costs billions of dollars. We’re going to be left with the OpenAI’s and Google’s of the world as the only players in the space until someone solves synthetic data.

Exactly this. I work at a small web scraping company (so I might be a bit bias) and any small business can collect a fair, capable datasets of public data for model training, sentiment analysis or whatever today. If public data is stopped by copyright as this lawsuit implies that would just mean only giant corporations and pirates would be able to afford this.

This would be a huge blow to open-source and research developers and I'd even argue it could help openAI to get a bit of a moat ala regulatory capture.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#664
post #645

I don't think the lawsuit has any merit, but I'd still like to encourage Sam Altman et al, if they really care about the greater good, to go Keyser Söze and immediately release torrents of the weights and source code for GPT-4 under GPL.

> I don't think the lawsuit has any merit The lawsuit fundamentally has merit. It asks a huge open question that no one knows the answer to. The outcome will be extraordinarily impactful. The question must be answered at some point. The case has merit even if NYT loses across the board.

I can see how someone can disagree with the NYT position but the idea that it lacks merit is wild!

AI might be the defining issue of copyright law for decades. There are so many open questions, and this seems like just the start.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#665
I asked an LLM to summarize the 69 page lawsuit. It does a decent job. Didn't infringe on any copyrights in the process :)

Here is a summary of the key points from the legal complaint filed by The New York Times against Microsoft and OpenAI:

The New York Times filed a copyright infringement lawsuit against Microsoft and OpenAI alleging that their generative AI tools like ChatGPT and Bing Chat infringe on The Times's intellectual property rights by copying and reproducing Times content without permission to train their AI models.

The Times invests enormous resources into producing high-quality, original journalism and has over 3 million registered copyrighted works. Its business models rely on subscriptions, advertising, licensing fees, and affiliate referrals, all of which require direct traffic to NYTimes.com.

The complaint alleges Microsoft and OpenAI copied millions of Times articles, investigations, reviews, and other content on a massive scale without permission to train their AI models. The models encode and "memorize" copies of Times works which can be retrieved verbatim. Defendants' tools like ChatGPT and Bing then display this protected content publicly.

OpenAI promised to freely share its AI research when founded in 2015 but pivoted to a for-profit model in 2019. Microsoft invested billions into OpenAI and provides all its cloud computing. Their partnership built special systems to scrape and store training data sets with Times content emphasized.

The complaint includes many examples of the AI models reciting verbatim excerpts of Times articles, showing they were trained on this data. It also shows the models fabricating quotes and attributing them to the Times.

Microsoft's integration of the OpenAI models into Bing Chat and other products boosted its revenues and market value tremendously. OpenAI's release of ChatGPT also made it hugely valuable. But their commercial success relies significantly on unlicensed use of Times works.

The Times attempted to negotiate a deal with Microsoft and OpenAI but failed, hence this lawsuit. Generating substitute products that compete with inputs used to train models does not qualify as "fair use" exemptions to copyright. The Times seeks damages and injunctive relief.

In summary, The New York Times alleges Microsoft and OpenAI's AI products infringe Times copyrights on a massive scale to unfairly benefit at The Times's expense. The Times invested heavily in content creation and controls how its work is used commercially. Using Times content without payment or permission to build competitive tools violates its rights under copyright law.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#666

Earlier quoted context omitted.

> This is a strong claim that just downloading articles into training data is what violates the copyright. That GTP outputs verbatim copies is a red herring. It's the other way around. There is no infringement if the model output is not substantially similar to a work in the training set [1]: > To win a claim of copyright infringement in civil or criminal court, a plaintiff must show he or she owns a valid copyright,…

The level of copying here is the copying into the training set , not the copying through use of the model . Its true that OpenAI will defend the wholesale copying into the training set by arguing that the transformative purpose of the next use reaches back and renders that copying fair use, but while that's clearly the dominant position of the AI industry, and it definitely seems compatible with the Cobstitutional pu…

> The level of copying here is the copying into the training set, not the copying through use of the model.

NY Times is suing because of both the model outputs and the existence of the training set. But infringement in the training set doesn't necessarily mean that the model infringes. Why? Because of the substantial similarity requirement. But first, I'll address the training set.

For articles that a person obtains through legal methods (like buying subscriptions) but doesn't then republish, storing copies of those articles is analogous to recording a legally accessed television show (time-shifting), which generally is fair use. Currently, no court has ruled that "analogous to time-shifting" is good enough for the time-shifting precedent to apply, but I think the difference is not significant. The same applies to companies. Companies are not literally people, but there isn't a reason for the time-shifting precedent to not apply to companies.

What about the articles that OpenAI obtained through illegal methods? Then the very act of obtaining those articles would be illegal. The training set contains those copies, so NY Times can sue to make OpenAI delete those copies and pay damages. But it's not trivially obvious that a GPT model is a copy of any works or contains copied expression of the any works in the training set; the weights that make up the model represent millions of works, it's not trivially obvious that the model contains something substantially similar to the expression in a work in the training set. Therefore, it's not trivially obvious that infringement with respect to the training set amounts to infringement with respect to the model made from the training set. If OpenAI obtained NY Times articles through illegal means, then making OpenAI delete the training set would be reasonable, but the model is a separate matter.

As long as the model doesn't contain copied expression and the weights can't be reversed into something substantially similar to expression in the existing works, then what matters is the output of the model.

If a user gives a prompt which contains no reference to an existing NY Times author, work, or a strongly associated characteristic/style, then do OpenAI's models produce outputs substantially similar to expression in the existing works? If not, then OpenAI shouldn't be liable for infringing works, because the infringing works result from the user's prompts. If my premise is false, then my conclusion falls apart. But if my premise is true, then at most I would admit that OpenAI has a limited burden to prevent users from giving those prompts.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#668

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

This analogy fails to capture the transformative nature of these models. Hosting a derivative work that is also a news article is not transformative. Hosting a next word completer is very different than a news article and can't be used as a substitute.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#669
post #418

People who think the examples the lawsuit are “fair use” need to consider what that would mean. We’re basically going to let a few companies consolidate all the value on the Internet into their black boxes with basically no rules … that seems very dangerous to me. I hope a court establishes some rules of engagement here, even if it’s not this case.

I see the exact opposite - any open source model is going to become prohibitively expensive to train if quality data costs billions of dollars. We’re going to be left with the OpenAI’s and Google’s of the world as the only players in the space until someone solves synthetic data.

This feels like a 1996 "music is too expensive for kids so they HAVE to pirate it."

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#670

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

What you described is entirely fair use, actually. Not only that, look at a few news articles from Tier 2 and down publications, and you'll realize that almost all of them are directly sourced from NYT and others. They'll say "so and so happened, according to The Times" (and usually link the article there)

it's fair use if you don't make money from your project no?
Post reply on HN