Live data from Hacker News

Anthropic agrees to pay $1.5B to settle lawsuit with book authors

nytimes.com

321–330 of 761 posts

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#321
post #8

To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.

> It’s important in the fair use assessment to understand that the training itself is fair use, I think that this is a distinction many people miss. If you take all the works of Shakespeare, and reduce it to tokens and vectors is it Shakespeare or is it factual information about Shakespeare? It is the latter, and as much as organizations like the MLB might want to be able to copyright a fact you simply cannot do that…

The question is going to be how much human intellectual input there was I think. I don't think it will take much - you can write the crappiest novel on earth that is complete random drivel and you still have copyright on it.

So to me, if you are doing literally any human review, edits, control over the AI then I think you'll retain copyright. There may be a risk that if somebody can show that they could produce exactly the same thing from a generic prompt with no interaction then you may be in trouble, but let's face it should you have copyright at that point?

This is, however, why I favor stopping slightly short of full agentic development at this point. I want the human watching each step and an audit trail of the human interaction in doing it. Sure I might only get to 5x development speed instead of 10x or 20x but that is already such an enormous step up from where we were a year ago that I am quite OK with that for now.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#322

Earlier quoted context omitted.

Afaik to scan a book you need to destroy it by cutting the spine so it can feed cleanly into the scanner. Would incur a lot of fines.

That's what they did. They also destroyed books worth millions in the process. They didn't think it would be a good idea to re-bind them and distribute it to the library or someone in need.

To be clear, they destructively scanned millions of books which in total were worth millions of dollars.

They did not destroy old, valuable books which individually were worth millions.

https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#323
post #270

Earlier quoted context omitted.

Option 1: $183B valuation, $1.5B settlement. Option 2: near-$0 valuation, $15M purchasing cost. To an investor, that just looks like a pretty good deal, I reckon. It's just the cost of doing business - which in my opionion is exactly what is wrong with practices like these.

> which in my opionion is exactly what is wrong with practices like these. What's actually wrong with this? They paid $1.5B for a bunch of pirated books. Seems like a fair price to me, but what do I know. The settlement should reflect society's belief of the cost or deterrent, I'm not sure which (maybe both). This might be controversial, but I think a free society needs to let people break the rules if they are willi…

I agree to some extent, but there is a slippery slope to “no rules apply to the rich”.

I do agree that in the case of victimless crimes, having some ability to recompensate for damages instead of outright ban the thing, means that we can enact many massively net-positive scenarios.

Of course, most crimes aren’t victimless and that’s where the negative reactions are coming from (eg company pollutes the commons to extract a profit).

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#324

Earlier quoted context omitted.

The settlement was for downloading the pirated books, not training from them. Unless they're paywalled it would be hard to argue the same for a blog.

It seems weird that there was legal culpability for downloading pirated books but not for training on them. At the very least, there is a transitive dependency between the two acts. Other people have said that Anthropic bought the books later on, but I haven't found any official records for that. Where would I find that? Also, does anyone know which Anthropic models were NOT trained on the pirated books. I want to av…

As far as anyone knows, no models were trained on the illegally downloaded books.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#325
post #315

Earlier quoted context omitted.

> which in my opionion is exactly what is wrong with practices like these. What's actually wrong with this? They paid $1.5B for a bunch of pirated books. Seems like a fair price to me, but what do I know. The settlement should reflect society's belief of the cost or deterrent, I'm not sure which (maybe both). This might be controversial, but I think a free society needs to let people break the rules if they are willi…

> I think a free society needs to let people break the rules if they are willing to pay the cost so you don't think super rich people should be bound by laws at all? Unless you made the cost proportional to (maybe expontial to) somebody's wealth, you would be creating a completely lawless class who would wreak havoc on society.

Hate to break it to you, but that's currently the world we live in. And yes, it sucks.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#326
post #270

Earlier quoted context omitted.

Option 1: $183B valuation, $1.5B settlement. Option 2: near-$0 valuation, $15M purchasing cost. To an investor, that just looks like a pretty good deal, I reckon. It's just the cost of doing business - which in my opionion is exactly what is wrong with practices like these.

> which in my opionion is exactly what is wrong with practices like these. What's actually wrong with this? They paid $1.5B for a bunch of pirated books. Seems like a fair price to me, but what do I know. The settlement should reflect society's belief of the cost or deterrent, I'm not sure which (maybe both). This might be controversial, but I think a free society needs to let people break the rules if they are willi…

[deleted]

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#327
post #310
post #308

Earlier quoted context omitted.

The lawsuit didn't find anything, Anthropic claimed this as part of the settlement. Companies settle without admission of wrongdoing all the time, to the extent that it can be bargained for.

The judge's ruling from earlier certainly seemed to me to suggest that the training was fair use. Obviously, that's not part of the current settlement. I'm no expert on this, so I don't know the extent to which the earlier ruling applies.

If I'm reading this right yes the training was fair use, but I was responding (unclearly) to the claim that the pirated books weren't used to train commercially released LLMs. The judge complained that it wasn't clear what was actually used, from the June order https://fingfx.thomsonreuters.com/gfx/legaldocs/jnvwbgqlzpw/... [pdf]:

> Notably, in its motion, Anthropic argues that pirating initial copies of Authors’ books and millions of other books was justified because all those copies were at least reasonably necessary for training LLMs — and yet Anthropic has resisted putting into the record what copies or even sets of copies were in fact used for training LLMs.

> We know that Anthropic has more information about what it in fact copied for training LLMs (or not). Anthropic earlier produced a spreadsheet that showed the composition of various data mixes used for training various LLMs — yet it clawed back that spreadsheet in April. A discovery dispute regarding that spreadsheet remains pending.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#328
Everything talks about settlement to the 'authors'; is that meant to be shorthand for copyright holders? Because there are a lot of academic works in that library where the publisher holds exclusive copyright and the author holds nothing.

By extension, if the big publishers are getting $3000 per article, that could be a fairly significant windfall.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#329
post #8

To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.

I think the jury is still out on how fair use applies to AI. Fair use was not designed for what we have now.

I could read a book, but its highly unlikely I could regurgitate it, much less months or years later. An LLM, however, can. While we can say "training is like reading", its also not like reading at all due to permanent perfect recall.

Not only does an LLM have perfect recall, it also has the ability to distribute plagiarized ideas at a scale no human can. There's a lot of questions to be answered about where fair use starts/ends for these LLM products.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#330
In related thought, when I listen to Suno, when I create "Epic Power Metal", the singer is very-often indistinguishible from the famous Hansi Kursch, of Blind Guardian.

https://en.wikipedia.org/wiki/Hansi_K%C3%BCrsch

I'm not sure if he even knows, but that is almost certainly his tracks they trained on.

Post reply on HN