Live data from Hacker News

Anthropic agrees to pay $1.5B to settle lawsuit with book authors

nytimes.com

361–370 of 761 posts

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#361
post #315

Earlier quoted context omitted.

> I think a free society needs to let people break the rules if they are willing to pay the cost so you don't think super rich people should be bound by laws at all? Unless you made the cost proportional to (maybe expontial to) somebody's wealth, you would be creating a completely lawless class who would wreak havoc on society.

Hate to break it to you, but that's currently the world we live in. And yes, it sucks.

I'm not sure how you're breaking that to me - it's the entire context of this discussion

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#362

Earlier quoted context omitted.

Yeah but did he die before anybody actually knew about it?

I knew about library genesis by 2012. It was at least 10 TiB large by then, IIRC. With the amount of Russian language content I got the impression it was more popular in that sphere, but an impressive collection for anyone and not especially secret.

To be fair, he might have been rather preoccupied at that time.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#363

Earlier quoted context omitted.

So how did they profit off the pirated books?

According to the judge, they didn't. The judge said they stored those books in a general purpose library for future use just in case they decided to use them later. It appears the judge took much issue with the downloading of "pirated content." And Anthropic decided to settle rather than let it all play out more.

But how the settlement cost was then defined if nobody read those books and there was no financial lost...

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#364
post #228

Earlier quoted context omitted.

> Buying used copies of books, scanning them, and training on it is fine. But nobody was ever going to that, not when there are billions in VC dollars at stake for whoever moves fastest. Everybody will simply risk the fine, which tends to not be anywhere close to enough to have a deterrent effect in the future. That is like saying Uber would have not had any problems if they just entered into a licensing contract wit…

Anthropic literally did exactly this to train its models according to the lawsuit. The lawsuit found that Anthropic didn't even use the pirated books to train its model. So there is that

I'm "team Anthropic" if we're stack ranking the major American labs pumping out SOTA models by ethics or whatever, but there is no universe in which a company like them operating in this competitive environment didn't pirate the books.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#365

If you are an author here are a couple of relevant links: You can search LibGen by author to see if your work is included. I believe this would make you a member of the class: https://www.theatlantic.com/technology/archive/2025/03/searc... If you are a member of the class (or think you are) you can submit your contact information to the plaintiff's attorneys here: https://www.anthropiccopyrightsettlement.com/

Thank you for posting this! I suspected my work was in the dataset and it looks like it is! I reached out via the form.

Good luck! Hope you get a payout!

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#366

Earlier quoted context omitted.

I think the jury is still out on how fair use applies to AI. Fair use was not designed for what we have now. I could read a book, but its highly unlikely I could regurgitate it, much less months or years later. An LLM, however, can. While we can say "training is like reading", its also not like reading at all due to permanent perfect recall. Not only does an LLM have perfect recall, it also has the ability to distrib…

> Not only does an LLM have perfect recall This has not been my experience. These days they are pretty good at googling though.

They do not have perfect recall unless you provide them a passage in the current context and then ask them to quote it.

The 'lossy encyclopedia' analogy is quite apt

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#368
post #8

To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.

> It’s important in the fair use assessment to understand that the training itself is fair use

IIUC this is very far from settled, at least in US law.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#369
post #277
post #228

Earlier quoted context omitted.

> Buying used copies of books, scanning them, and training on it is fine. But nobody was ever going to that, not when there are billions in VC dollars at stake for whoever moves fastest. Everybody will simply risk the fine, which tends to not be anywhere close to enough to have a deterrent effect in the future. That is like saying Uber would have not had any problems if they just entered into a licensing contract wit…

> But nobody was ever going to that Didn't Google have a long standing project to do just that? https://en.wikipedia.org/wiki/Google_Books

Crazy to think we've been helping train AI through captchas long before the "click all squares containing" ones.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#370
post #8

To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.

> Buying used copies of books, scanning them, and training on it is fine.

Awesome, so I just need enough perceptrons to overfit every possible copyrighted works then?

Post reply on HN