Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

671–680 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#671

I asked an LLM to summarize the 69 page lawsuit. It does a decent job. Didn't infringe on any copyrights in the process :) Here is a summary of the key points from the legal complaint filed by The New York Times against Microsoft and OpenAI: The New York Times filed a copyright infringement lawsuit against Microsoft and OpenAI alleging that their generative AI tools like ChatGPT and Bing Chat infringe on The Times's…

A Second LLM's take on this lawsuit can be found below. I'd love to see OpenAI address these complaints publicly and without incurring any additional damages to NYT.

The document is a legal complaint filed by The New York Times Company against Microsoft Corporation and various OpenAI entities, alleging copyright infringement and other related claims. The New York Times Company (The Times) accuses the defendants of unlawfully using its copyrighted works to create artificial intelligence (AI) products that compete with The Times, particularly generative artificial intelligence (GenAI) tools and large language models (LLMs). These tools, such as Microsoft's Bing Chat and OpenAI's ChatGPT, allegedly copy, use, and rely heavily on The Times’s content without permission or compensation.

Nature of the Action: The Times emphasizes the importance of independent journalism to democracy and claims its ability to continue providing this service is threatened by the defendants' actions. The complaint argues that the GenAI tools are built upon unlawfully copied New York Times content, which undermines The Times's investments in journalism.

Defendants: The defendants include Microsoft Corporation and various OpenAI entities, such as OpenAI Inc., OpenAI LP, and several other related companies. The Times alleges these entities have worked together to create and profit from the GenAI tools in question.

Allegations: 1. Copyright Infringement: The Times claims the defendants copied millions of its copyrighted articles and other content to train their GenAI models. This training allegedly involves large-scale copying and use of The Times’s content, emphasizing its quality and value in building effective AI models.

2. Unlawful Competition: The Times argues that the defendants' GenAI tools compete with it by providing access to its content for free, which could potentially divert readers and revenue away from The Times.

3. Misattribution and Hallucinations: The Times asserts that the defendants' tools not only unlawfully distribute its content but also generate and attribute false information to The Times, damaging its credibility and trust with readers.

4. Trademark Dilution: The complaint includes claims that the defendants' use of The Times’s trademarks in connection with lower-quality or inaccurate AI-generated content dilutes and tarnishes its brand.

5. Digital Millennium Copyright Act Violations: The Times alleges that the defendants removed or altered copyright management information from its works, which is prohibited under the law.

Harm to The Times: The Times claims it has suffered significant harm from these actions, including loss of control over its content, damage to its reputation for accuracy and quality, and financial losses due to diminished traffic and revenue.

Demands: The Times seeks various forms of relief, including statutory damages, injunctive relief to prevent further infringement, destruction of the infringing AI models, and compensation for losses and legal fees.

Overall Summary: This legal complaint represents a significant clash between traditional media and emerging AI technology companies. It underscores the complex legal, ethical, and economic issues arising from the use of copyrighted content to train AI systems. The outcome of this case could have far-reaching implications for the AI industry, content creators, and the broader digital ecosystem.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#672

Earlier quoted context omitted.

It would be curious to test this on a larger sample than just a few. It is hard to believe that a majority of NYT articles are verbatim stored in the weights of a web-wide LLM, but if that is the case it would be a pretty unbelievable revelation about their ability to compress an entire web’s worth of data. But, more likely, I assume it is a case of overfitting, or simply finding a prompt that happened to work well.…

You can't reproduce on the web interface, because the temperature settings are higher than what's required to compress the text. You need to use the API. However, I had good luck reproducing poems on GPT 3.5, both copyrighted and not copyrighted, because the choice of words is a lot more "specific" so to speak, and therefore higher temperature isn't enough to prevent complete reproduction of the originals. See https:…

It doesn’t seem that surprising; compared to entire NYT articles, poems are short, structured and more likely to be shared in multiple places across the web.

I’m more surprised that it can repeat 100 articles; if that behaviour is consistent in larger sample sizes and beyond just NYT dataset (which might be repeated on the web more than other sources, causing overfitting), that would be impressive.

You could imagine at some point a large enough GPT5 or 6 or 7 will be able to memorize verbatim every corner of the web.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#673
post #670

Earlier quoted context omitted.

What you described is entirely fair use, actually. Not only that, look at a few news articles from Tier 2 and down publications, and you'll realize that almost all of them are directly sourced from NYT and others. They'll say "so and so happened, according to The Times" (and usually link the article there)

it's fair use if you don't make money from your project no?

No.

In the US, whether or not you make money has little to do with whether or not your use qualifies as "fair use".

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#674

Earlier quoted context omitted.

My impression is that it’s not necessarily legal, but going after bloggers and proving damages based is just a huge waste of their time. OpenAI came by with their fat stack of funding and changed that.

No, in US law at least there can be no copyright of facts, only presentation. If you convey the same facts in different words that isn't a matter of fair use, it's never even a matter of copyright in the first place.

How about things that aren’t quite facts? Reviews, opinions, etc.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#675
post #492

This, or a lawsuit like it is going to be the SCO vs IBM of the 2020's, to wit: a copyright troll trying to extract rent, with various special interests cheering it on to try and promote their own agenda (ironically it was Microsoft that played that role with SCO). It's funny how times have changed and at least now a louder group seem to be on the troll's side. I hope to see some better analysis on the frivolity of t…

That is an irrelevant comparison.

This is theft and monstrous profit from theft. For actual justice this should be a class action suit of the world vs. OpenAI/Microsoft and the financial consequences should be company-ending for OpenAI. Otherwise, you have incented everyone in the AI industry to steal as much as they can for as long as they can.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#677
post #645

I don't think the lawsuit has any merit, but I'd still like to encourage Sam Altman et al, if they really care about the greater good, to go Keyser Söze and immediately release torrents of the weights and source code for GPT-4 under GPL.

> I don't think the lawsuit has any merit The lawsuit fundamentally has merit. It asks a huge open question that no one knows the answer to. The outcome will be extraordinarily impactful. The question must be answered at some point. The case has merit even if NYT loses across the board.

How is the question being asked different from the Google Books case?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#678
post #493

Earlier quoted context omitted.

> sure, but if I use an LLM to write a novel/article, I can be sued in civil court not the LLM That's function of the legal system, not of the technology. If tomorrow someone made a perfect dolphin-Esperanto translator and proved Dolphins were as smart as humans, you still can't sue a dolphin until the legal system says so.

Wouldn't you find out by suing the dolphin and seeing if it holds up in court?

Not if you were smart, unless you have some sort of solid argument for why the established case law about this sort of thing is faulty.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#679

The NYT is preparing for a tsunami by building a sandcastle. Big picture, this suit won’t matter, for so many reasons. To enumerate a few: 1. Next gen LLMs will be trained exclusively on “synthetic”/public data. GPT-4V can easily whitewash its entire copyrighted training corpus to be unrecognizably distinct (say reworded by 40%, authors/sources stripped, etc). Ergo there will be no copyright material for GPT-5 to reg…

> 2. Research/hosting/progress will proceed. The US cannot stop this, only choose to be left behind. The world will move on, with China gleefully watching as their biggest rival commits intellectual suicide all to appease rent seeking media companies.

Sorry, is this the same China that has already introduced their own sweeping regulations on AI? Which in at least one instance forced a Chinese startup to shut down their newly launched chatbot because it said things that didn't align with the party's official stance on the war in Ukraine?

https://finance.yahoo.com/news/beijing-tries-regulate-china-...

https://nitter.unixfox.eu/CDT/status/1625936306814717952?337...

I don't disagree that research/hosting/progress will continue, but I'm not so sure that it's China who stands to benefit from the US adding some guardrails to this rollercoaster.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#680
post #677

Earlier quoted context omitted.

> I don't think the lawsuit has any merit The lawsuit fundamentally has merit. It asks a huge open question that no one knows the answer to. The outcome will be extraordinarily impactful. The question must be answered at some point. The case has merit even if NYT loses across the board.

How is the question being asked different from the Google Books case?

You’re not just display the contents of copyrighted works publicly, they’re selling access. This flips the script for the 1st factor of the Fair Use test. Additionally, by selling it to people who use it to get news summaries, you can argue you damage the market for a NY Times subscription, which triggers the 4th factor.
Post reply on HN