Live data from Hacker News

Anthropic agrees to pay $1.5B to settle lawsuit with book authors

nytimes.com

251–260 of 761 posts

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#251
post #188

Earlier quoted context omitted.

I wonder what Aaron Swartz would think if he lived to see the era of libgen.

Didn't he get in trouble for contributing to sci-hub before he died?

He got into trouble for breaking into an unsecured network closet at MIT and using MIT credentials to download a bunch of copyrighted content.

The whole incident is written up in detail, https://swartz-report.mit.edu/ by Hal Abelson (who wrote SICP among other things). It is a well-researched document.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#252
post #18

Earlier quoted context omitted.

It is related to scalable mode training, however. Chopping the spine off books and putting the pages in an automated scanner is not scalable. And don't forget about the cost of 1) finding 2) purchasing 3) processing and 4) recycling that volume of books.

> Chopping the spine off books and putting the pages in an automated scanner is not scalable. That's how Google Books, the Internet Archive, and Amazon (their book preview feature) operated before ebooks were common. It's not scalable-in-a-garage but perfectly scalable for a commercial operation.

I don't think Google Books scanner chopped off the spine. https://linearbookscanner.org/ is the open design they released.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#253
post #141

Earlier quoted context omitted.

Yes, much. And they actually went and did that afterwards. They just pirated them first.

Where can I find source that says Anthropic bought the pirated books afterwards? I haven't seen this in any official document. Also, do we know if the newer models were trained without the pirated books?

> Where can I find source that says Anthropic bought the pirated books afterwards? I haven't seen this in any official document.

https://storage.courtlistener.com/recap/gov.uscourts.cand.43...

> Also, do we know if the newer models were trained without the pirated books?

I'm pretty sure we do but I couldn't swear to it or quickly locate a source.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#254

Settlement Terms (from the case pdf) 1. A Settlement Fund of at least $1.5 Billion: Anthropic has agreed to pay a minimum of $1.5 billion into a non-reversionary fund for the class members. With an estimated 500,000 copyrighted works in the class, this would amount to an approximate gross payment of $3,000 per work. If the final list of works exceeds 500,000, Anthropic will add $3,000 for each additional work. 2. Des…

Don't forget: NO LEGAL PRECEDENT! which means, anybody suing has to start all over. You only settle in this scenario/point if you think you'll lose. Edit: I'll get ratio'd for this- but its the exact same thing google did in it's lawsuit with Epic. They delayed while the public and courts focused in apple (oohh, EVIL apple)- apple lost, and google settled at a disadvantage before they had a legal judgment that couldn…

A full case is many more years of suits and appeals with high risks, so its natural to settle which obviously means no precedent

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#255

This is sad for open source AI, piracy for the purpose of model training should also be fair use because otherwise only the big companies who can afford to pay off publishers like Anthropic will be able to do so. There is no way to buy billions of books just for model training, it simply can't happen.

This is a settlement. It does not set a precedent nor even admit to wrongdoing.

> otherwise only the big companies who can afford to pay off publishers like Anthropic will be able to do so

Only well funded companies can afford to hire a lot of expensive engineers and train AI models on hundreds of thousands of expensive GPUs, too.

Something tells me many the grassroots LLM training people are less concerned about legality of their source training set than the big companies anyway.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#256

Earlier quoted context omitted.

It implies that people want everyone to do this when it's clear no one should do it. I'm not exactly a fan of "this isn't profitable for small businesses to steal from so we should make it so everyone should steal".

Piracy is not stealing. I don't know why everyone on HN suddenly turned into a copyright hawk, only big companies benefit from our current copyright regime, like Disney and their lobbying for increasing its length.

> only big companies benefit from our current copyright regime

You’ve never authored, created, or published something? Never worked for a company that sells something protected by copyright?

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#257

Earlier quoted context omitted.

I didn't realize Vernor Vinge had passed away... Sad TIL

There was a nice discussion & nostalgia at the time (1151 points, 2024, 320 comments) https://news.ycombinator.com/item?id=39775304

Cookie monster is his strongest work. It has a VIBE.

Reminds me of permutation city

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#258
post #145
post #67

One thing that comes to mind is... Is there a way to make your content on the web "licensed" in a way where it is only free for human consumption? I.e. effectively making the use of AI crawlers pirating, thus subject to the same kind of penalties here?

Yes to the first part. Put your site behind a login wall that requires users to sign a contract to that effect before serving them the content... get a lawyer to write that contract. Don't rely on copyright. I'm not sure to what extent you can specify damages like these in a contract, ask the lawyer who is writing it.

Contracts generally require an exchange of consideration (something of value, like money).

If you put a “contract” on your website that users click through without paying you or exchanging value with you and then you try to collect damages from them according to your contract, it’s not going to get you anywhere.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#259

So if a startup wants to buy book PDFs legally to use for AI purposes, any suggestions on how to do that?

Reach the publishers or resellers (like amazon for instance)

Give them this order : "I want to buy all your books as epub"

Pay and fetch the stuff

That's all

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#260
post #145

Earlier quoted context omitted.

Yes to the first part. Put your site behind a login wall that requires users to sign a contract to that effect before serving them the content... get a lawyer to write that contract. Don't rely on copyright. I'm not sure to what extent you can specify damages like these in a contract, ask the lawyer who is writing it.

Contracts generally require an exchange of consideration (something of value, like money). If you put a “contract” on your website that users click through without paying you or exchanging value with you and then you try to collect damages from them according to your contract, it’s not going to get you anywhere.

The consideration the viewer received was access to your private documents.

The consideration you received was a promise to refrain from using those documents to train AI.

I'm not a lawyer, but by my understanding of contract law consideration is trivially fulfilled here.

Post reply on HN