Live data from Hacker News

Anthropic agrees to pay $1.5B to settle lawsuit with book authors

nytimes.com

391–400 of 761 posts

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#392
post #8

To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.

> It’s important in the fair use assessment to understand that the training itself is fair use IIUC this is very far from settled, at least in US law.

Yes, but if you are predisposed for some reason to think that Anthropic "won" this case, then you're going to believe all sorts of things.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#393
post #386

From a systems design perspective, $3,000 per book makes this approach completely unscalable compared to web scraping. It's like choosing between a O(n) and O(n²) algorithm - legally compliant data acquisition has fundamentally different scaling characteristics than the 'move fast and break things' approach most labs took initially.

more of a large difference in constant factor, like a galactic algorithm for data trawling

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#395

So the article notes Anthropic states they never publicly released a frontier model that was trained on the downloaded copyright material. So were Claude 2 and 3 only trained on legally purchased and scanned books, or do they now use a different training system that does not rely on books at all ?

it sounds like the former

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#396
post #263

Why are they paying $3000 per book. Does anyone think these authors srll their books for that amount?

If you acquire something illegally of course the judgement against you has to be much higher than the legal price. Why would anyone purchase anything if the worst thing that could happen to you for stealing it was just paying the retail price?

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#397
post #353
post #335

Earlier quoted context omitted.

> Since when? Since in our legal system, only humans and groups of humans (the corporation is a convenient legal proxy for a group of humans that have entered into an agreement) have rights. Property doesn't have rights. Land doesn't have rights. Books don't have rights. My computer doesn't have rights. And neither does an LLM.

Maybe we should give machines rights, then.

Maybe we should. Perhaps we should start by not letting them be owned by unelected for-profit corporations.

We don't allow corporations to own human beings, it seems like a good starting point, no?

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#398
post #386

From a systems design perspective, $3,000 per book makes this approach completely unscalable compared to web scraping. It's like choosing between a O(n) and O(n²) algorithm - legally compliant data acquisition has fundamentally different scaling characteristics than the 'move fast and break things' approach most labs took initially.

I don't know if anyone has actually read the article or the ruling, but this is about pirating books.

Anthropic went back and bought->scanned->destroyed physical copies of them afterward... but they pirated them first, and that's what this settlement is about.

The judge also said:

> “The training use was a fair use,” he wrote. “The technology at issue was among the most transformative many of us will see in our lifetimes.”

So you don't need to pay $3,000 per book you train on unless you pirate them.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#399
post #262

Earlier quoted context omitted.

Sure, but that’s mostly because the sheer convenience of the illegal way is so much higher, and carries zero startup cost.

The same could be said of grand larceny. The difference would seem to be a mix of social norms and, more notably for this conversation, very different consequences.

Not sure it is realistic or easier to physically steal 500k books.

I get what you are going for, but my point was that a dataset existed, and the only way it could be compiled was illegaly.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#400
post #228

Earlier quoted context omitted.

> Buying used copies of books, scanning them, and training on it is fine. But nobody was ever going to that, not when there are billions in VC dollars at stake for whoever moves fastest. Everybody will simply risk the fine, which tends to not be anywhere close to enough to have a deterrent effect in the future. That is like saying Uber would have not had any problems if they just entered into a licensing contract wit…

> It was faster to just put unlicensed taxis on the streets and use investor money to pay fines and lobby for favorable legislation And thank god they did. There was no perfectly legal channel to fix the taxi cartel. Now you don't even have to use Uber in many of these places because taxis had to compete - they otherwise never would have stopped pulling the "credit card reader is broken" scam, taking long routes on p…

i dont know that its such a great thing in the end. Uber/Lyft is 50-100% more expensive now than taxis were before. Theyre entrenched in different ways.
Post reply on HN