Live data from Hacker News

Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

apnews.com

521–530 of 654 posts

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#521

Earlier quoted context omitted.

> similarly, as has been shown (by numerous publishers and authors), LLMs are capable of producing nearly-verbatim versions of the texts they have been trained on, to a well enough quality that most people won't be able to notice the difference. If that is true, you have a legal claim and can sue them. I doubt that’s true in the general case though. The “does it hurt the original publisher” is a test for fair use BTW…

It's a perfectly sensible interpretation of international copyright law. I'm not sure if you're serious with the suggestion I could sue them. These are both US corporations, that justice system is pretty much in shambles in particular when it concerns corporations as big as these AI ones. You can dig your heels in the sand to defend that system, but you will also have to dig your head in the sand about why Sam Altman…

> Also even if it would work (which it very much doesn't), whether it "hurts the original publisher" is actually literally one of the criteria for considering something fair use or not. Look it up.

I don’t know why you’re repeating the stuff I just wrote like I didn’t. My point is that this is only relevant for the fair use defense and not copyright in general.

Here’s what I said:

> The “does it hurt the original publisher” is a test for fair use BTW, just because you hurt the sales of someone doesn’t necessarily make it copyright infringement. That is only relevant if you try to defend using fair use (and it’s only part of the test that’s used to decide fair use).

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#522
post #239

Earlier quoted context omitted.

So, they can bu ya book and format shift it, but when I do it, it's piracy? Looking at all DMCA/DRM systems.

Format shifting has a DMCA carveout. It is 100% allowed. Since around 2000, the rule has permitted it for: > Literary works, including computer programs and databases, protected by access control mechanisms that fail to permit access because of malfunction, damage, or obsoleteness. DRM being covered under other laws, and being gross, still applies. And still applies to industry giants, too. Which is why most who do t…

Yes, it's legal to format-shift DRM protected media, but it is not lawful. Someone has to break the law to provide me with a decryption tool.

If you asked the right politician when all these rules were being written, the intent was that each person who needs to format-shift their media would independently write their own decryption tools, use them for lawful purposes only, and then dutifully delete them the moment they were no longer needed. This is, of course, laughable.

Of course, if Anthropic was, say, buying and decrypting Kindle books TODAY; they probably could get Claude to vibe-code a DRM decryption tool[0]. That would actually be within the bounds of this asinine law. If Anthropic started off by doing this, however, they probably would have just used a decryption tool found on the Internet, and that would have invited different legal challenges. Like, is it legal to use an unlawful tool to accomplish something otherwise legally protected? The courts so far have been very hostile to ANY attempt to tie the anticircumvention provisions of the DMCA to fair use. They could easily say "No, you only get to format shift with your own tools".

[0] Related note: I really wish I had Mythos access, just so I could jailbreak my iPad on modern iPadOS. No other reason.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#523
post #453

Earlier quoted context omitted.

Whatever "training" is, if you can't persuade the machine to spit substantially the same text back out verbatim, it's clearly not something that falls under copy right law either, because there's no copy. Yes, for some texts that's possible. But for the vast majority, it is not.

Can you cite any information on this not being possible for the vast majority? Or is it simply that the correct prompt hasn't been written for all possible cases? I also fail to see the difference if logic/harnessing is added around a vector database that can output the complete corpus, but simply is instructed not to. It very clearly is still compressing the information into the vector weights, and then recovering t…

Nah, I can't prove a negative. But Common Crawl is 12 petabytes and is not the largest part of what these models get trained on. DeepSeek v4 Pro is, what, 865GB?

That's one hell of a compression ratio, if it can do what you claim.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#524

Earlier quoted context omitted.

Information entropy. The amount of data an LLM ingests cannot be compressed to the size of the weights even at maximum theoretical compression.

Is that relevant? I can use a lossy compression algorithm such that the original could never be recovered from the image I've produced, but that derived image would surely be under copyright. LLMs are obviously capable of producing "exact" phrases as well. Ask it to give you famous quotes, it can do it. Ask it to read a paper for you and cite it, it can do it.

> but that derived image would surely be under copyright.

I wouldn't bet on that. https://en.wikipedia.org/wiki/Campbell%27s_Soup_Cans

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#525

Earlier quoted context omitted.

Can you cite any information on this not being possible for the vast majority? Or is it simply that the correct prompt hasn't been written for all possible cases? I also fail to see the difference if logic/harnessing is added around a vector database that can output the complete corpus, but simply is instructed not to. It very clearly is still compressing the information into the vector weights, and then recovering t…

You are asking to prove a negative. But even assuming that the model is capable of returning every bit of its training data verbatim (a mathematical impossibility) that would not be enough as mere capability is insufficient here. If capability alone were the standard any library that also has a photocopier / scanner would be in violation. To prove distribution of copyrighted materials it would have to be practical an…

There was a paper a while back where (from memory) they managed to coax 75% of the original text of some internet-popular books out of an LLM. Harry Potter, 1984, etc. That's why I said it was possible for some texts.

My assumption is that multiple copies in the training data "wear a deeper groove". I believe those are infringing, and should be dealt with on a case-by-case basis. But the vast majority of text doesn't wear that groove.

(Edit: Think it was this one https://arxiv.org/abs/2601.02671)

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#526
post #487

Earlier quoted context omitted.

As far as I'm concerned, the courts are wrong, and training on ill gotten copyrighted material is not fair use. Given the clear value of highly trained LLMs, the investment they have taken on, and the amount of disruption to the existing economy they stand to make, in a just world, the people who created the training data deserve some level of compensation. I think, in the US, they are very afraid of falling behind C…

IMO, using copyrighted works to train models should only be "fair use", if the models are then released as (at least) open weight, so that the public can benefit from it. (Although as noted by a sibling, this would require a law change, not action by the court).

I’m not sure if that’s enough but it would be a great start.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#527
post #493
post #453

Earlier quoted context omitted.

Whatever "training" is, if you can't persuade the machine to spit substantially the same text back out verbatim, it's clearly not something that falls under copy right law either, because there's no copy. Yes, for some texts that's possible. But for the vast majority, it is not.

> spitting out verbatim text The New York Times lawsuit is resting on the point that large chunks of undigested articles can be vomited out. OpenAI tried to have the lawsuit thrown out but the courts permitted it to continue. The Times... alleged that OpenAI's ChatGPT and Microsoft's Copilot had produced near-verbatim replicas of copyrighted articles, that the chatbots generated hallucinated content falsely attribute…

It's possible. Would be interesting to see their evidence, and to know whether they can reproduce it for arbitrary articles, not just ones that have been endlessly republished on the net.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#528
post #518
post #493

Earlier quoted context omitted.

> spitting out verbatim text The New York Times lawsuit is resting on the point that large chunks of undigested articles can be vomited out. OpenAI tried to have the lawsuit thrown out but the courts permitted it to continue. The Times... alleged that OpenAI's ChatGPT and Microsoft's Copilot had produced near-verbatim replicas of copyrighted articles, that the chatbots generated hallucinated content falsely attribute…

> spitting out verbatim text > had produced near-verbatim replicas

Don't get too hung up on the preciseness of the copy - the courts won't. I doubt that spitting out an existing article with a few adjectives changed would be considered transformative.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#529
post #524

Earlier quoted context omitted.

Is that relevant? I can use a lossy compression algorithm such that the original could never be recovered from the image I've produced, but that derived image would surely be under copyright. LLMs are obviously capable of producing "exact" phrases as well. Ask it to give you famous quotes, it can do it. Ask it to read a paper for you and cite it, it can do it.

> but that derived image would surely be under copyright. I wouldn't bet on that. https://en.wikipedia.org/wiki/Campbell%27s_Soup_Cans

I don't know that this really challenges anything relating to compression.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#530

A good example of the problem with this settlement: >It appears that LLMs have already incorporated APOSD EDIT: The text of the book _A Philosophy of Software Design_ ENDEDIT (which would seem to be illegal, since it is copyrighted). For example, I have asked ChatGPT questions about APOSD and it seems to be able to answer. https://groups.google.com/g/software-design-book/c/_wl1DciZZ...

I don't see: - what's APOSD - "it's illegal since it's copyrighted" makes no sense to me - The settlement should be exactly to cover their licenses for training

Edited to clarify APOSD == the book _A Philosophy of Software Design_

Please ask John Ousterhout what his cut of this settlement will be, and whether or no he agreed to it and finds it acceptable.

Post reply on HN