Live data from Hacker News

Anthropic agrees to pay $1.5B to settle lawsuit with book authors

nytimes.com

621–630 of 761 posts

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#621

Earlier quoted context omitted.

It’s the sign of a health economy when we respect the creation of content.

It's a sign of rent seeking economy in decline. Rising economies never respect IPs.

IP protections are in the US Constitution. Has the US been in decline since the late 1700s?

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#622
post #441
post #248

I can't help but feel like this is a huge win for Chinese AI. Western companies are going to be limited in the amount of data they can collect and train on, and Chinese (or any foreign AI) is going to have access to much more and much better data.

I think western companies will be just fine -- Anthropic is settling because they illegally pirated books from LibGen back in 2021 and subsequently trained models on them. They realized this was an issue internally and pivoted to buying books en masse and scanning them into digital formats, destroying the original copies in the process (they actually hired a former lead in the Google Books project to help them in thi…

It's easier for one company to digitize and sell/share than for many companies to do it individually.

Western companies will be fine but sharing data in ways that would be illegal in the US does help other companies outside the US.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#623
post #386

From a systems design perspective, $3,000 per book makes this approach completely unscalable compared to web scraping. It's like choosing between a O(n) and O(n²) algorithm - legally compliant data acquisition has fundamentally different scaling characteristics than the 'move fast and break things' approach most labs took initially.

Isn't a flat price per book quite plainly O(n)? If not, what's n?

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#624

Earlier quoted context omitted.

i dont know that its such a great thing in the end. Uber/Lyft is 50-100% more expensive now than taxis were before. Theyre entrenched in different ways.

Uber did a great job convincing lay people that taxis were ripoffs and they were a good deal. For some time that was probably true. Now, I see people at the airport walk over to the pickup lot, joining a crowd of others furiously messing with their phones while scanning the area for presumably their driver. All the while the taxis waiting immediately outside the exit door were $2 more expensive, last time I checked.

Not at any airport I've been to recently. I've never seen lines of taxis waiting at any airport in the last few years. There are empty taxi slots. People hail the taxi using an app and then wait for it to show up. Just like Lyft/Uber.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#625
post #8

To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.

> the training itself is fair use

Sure, training by itself isn't worth anything.

Distributing and collecting payment for the usage of a trained model which may violate copyright, etc. that's still an open legal question and worth billions as well.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#626
post #320
post #311

Earlier quoted context omitted.

> Everyone has more than a right to freely have read everything is stored in a library. Every human has the right to read those books. And now, this is obvious, but it seems to be frequently missed - an LLM is not a human , and does not have such rights.

By US law, cccording to Author's Guild vs Google[1] on the Google book scanning project, scanning books for indexes is fair use. Additionally: > Every human has the right to read those books. Since when? I strongly disagree - knowledge should be free. I don't think the author's arrangement of the words should be free to reproduce (ie, I think some degree of copyright protection is ethical) but if I want to use a tool…

Scanning books for indexes is fair use. Very notably providing access to those books to the public for free was not fair use...

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#627
post #421

Settlement Terms (from the case pdf) 1. A Settlement Fund of at least $1.5 Billion: Anthropic has agreed to pay a minimum of $1.5 billion into a non-reversionary fund for the class members. With an estimated 500,000 copyrighted works in the class, this would amount to an approximate gross payment of $3,000 per work. If the final list of works exceeds 500,000, Anthropic will add $3,000 for each additional work. 2. Des…

So they can also keep models trained on the datasets? That seems pretty big too, unless the half life of models is so low it doesn't matter.

It's a separate suit being wages against Meta and OpenAI etc.

There's piracy, then there's making available a model to the public which can regurgitate copyrighted works or emulate them. The latter is still unsettled

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#628

Earlier quoted context omitted.

I think the jury is still out on how fair use applies to AI. Fair use was not designed for what we have now. I could read a book, but its highly unlikely I could regurgitate it, much less months or years later. An LLM, however, can. While we can say "training is like reading", its also not like reading at all due to permanent perfect recall. Not only does an LLM have perfect recall, it also has the ability to distrib…

> I could read a book, but its highly unlikely I could regurgitate it, much less months or years later. And even if one could, it would be illegal to do. Always found this argument for AI data laundering weird.

But there is a difference between “illegal to regurgitate it” and “illegal to remember it”. IIRC in this case that settled the judge had ruled on “remember” (fair use) but not on the other.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#629

Earlier quoted context omitted.

> It was faster to just put unlicensed taxis on the streets and use investor money to pay fines and lobby for favorable legislation And thank god they did. There was no perfectly legal channel to fix the taxi cartel. Now you don't even have to use Uber in many of these places because taxis had to compete - they otherwise never would have stopped pulling the "credit card reader is broken" scam, taking long routes on p…

i dont know that its such a great thing in the end. Uber/Lyft is 50-100% more expensive now than taxis were before. Theyre entrenched in different ways.

I strongly prefer to take traditional taxis, but I also comparison shop and Lyft is almost always 20-40% cheaper than a cab ride.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#630

Earlier quoted context omitted.

The law was not broken by "super rich people". It was broken by a company of people who were not very rich at all and have managed to produce billions in value (not dollars, value) by breaking said laws. They're not trafficking humans or doing predatory lending, they're building AI. This is why our judicial system literally handles things on a case by case basis.

I just want to make sure I understand this correctly. Your argument is that this is all fine because it wasn't done by people who were super rich but instead done by people who became super rich and were funded by the super rich? I just want to check that I have that right. You are arguing that if I'm a successful enough bank robber that this is fine because I pay some fine that is a small portion of what I heisted?…

You're asking me if straight up stealing money from a bank is comparable to stealing books in 2025 to train an AI which will generate untold value for people?
Post reply on HN