Live data from Hacker News

Anthropic agrees to pay $1.5B to settle lawsuit with book authors

nytimes.com

161–170 of 761 posts

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#161
post #107

Great. Which rich person is going to jail for breaking the law?

No one, rich or poor, goes to jail for downloading books.

Are you sure? I think in some jurisdictions they would, according to the law.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#162

Settlement Terms (from the case pdf) 1. A Settlement Fund of at least $1.5 Billion: Anthropic has agreed to pay a minimum of $1.5 billion into a non-reversionary fund for the class members. With an estimated 500,000 copyrighted works in the class, this would amount to an approximate gross payment of $3,000 per work. If the final list of works exceeds 500,000, Anthropic will add $3,000 for each additional work. 2. Des…

I’m an author, can I get in on this?

I had the same question.

It looks like you'll be able to search this site if the settlement is approved:

> https://www.anthropiccopyrightsettlement.com/

If your work is there, you qualify for a slice of the settlement. If not, you're outta luck.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#163
post #120

Earlier quoted context omitted.

If in most cases damages cannot be recovered or the criminal will never be caught in the first place, then what is the lesson being taught? Doesn't that just create a moral hazard where you "randomly" choose who to penalize?

It's about sending a message.

The message being you’ll likely get away with it?

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#164
post #8

To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.

> Buying used copies of books, scanning them, and training on it is fine.

Buying used copies of books, scanning them, and printing them and selling them: not fair use

Buying used copies of books, scanning them, and making merchandise and selling it: not fair use

The idea that training models is considered fair use just because you bought the work is naive. Fair use is not a law to leave open usage as long as it doesn’t fit a given description. It’s a law that specifically allows certain usages like criticism, comment, news reporting, teaching, scholarship, or research. Training AI models for purposes other than purely academic fits into none of these.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#165

Wait, I’m a published author, where’s my check

The court has to give preliminary approval to the settlement first. After that there should be a notice period during which the lawyers will attempt to reach out and tell you what you need to do to receive your money. (Not a lawyer, not legal advice).

You can follow the case here: https://www.courtlistener.com/docket/69058235/bartz-v-anthro...

You can see the motion for settlement (what the news article is about) here: https://storage.courtlistener.com/recap/gov.uscourts.cand.43...

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#167

This weirdly seems like its the best mechanism to buy this much data. Imagine going to 500k publishers to buy it individually. 3k per book is way cheaper. The copyright system is turning into a data marketplace in front of our eyes

I suspect you could acquire and scan every readily purchasable book for much less than $3k each. Scanhouse for instance charges $0.15 per page for regular unbound (disassembled) books, plus $0.25 for supervised OCR, plus another dollar if the formatting is especially complex; this comes out to maybe $200-300 for a typical book. Acquiring, shipping, and disposing of them all would of course cost more, but not thousands more.

The main cost of doing this would be the time - even if you bought up all the available scanning capacity it would probably take months. In the meantime your competition who just torrented everything would have more high-quality training data than you. There are probably also a fair number of books in libgen which are out of print and difficult to find used.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#168
This was a very tactical decision by Anthropic. They have just received Series F funding, and they can now afford to settle this lawsuit.

OpenAI and Google will follow soon now that the precedent has been set, and will likely pay more.

It will be a net win for Anthropic.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#170

Earlier quoted context omitted.

> It’s important in the fair use assessment to understand that the training itself is fair use, I think that this is a distinction many people miss. If you take all the works of Shakespeare, and reduce it to tokens and vectors is it Shakespeare or is it factual information about Shakespeare? It is the latter, and as much as organizations like the MLB might want to be able to copyright a fact you simply cannot do that…

> And what about all the other stuff that LLM's spit out? Who owns that. Well at present, no one. If you train a monkey or an elephant to paint, you cant copyright that work because they aren't human, and neither is an LLM. This seems too cute by half, courts are generally far more common sense than that in applying the law. This is like saying using `rails generate model:example` results in a bunch of code that isn'…

The example is a real legal case afaik, or perhaps paraphrased from one (don’t think it was a monkey - an ape? An elephant?).

I’d guess the legal scenario for `rails generate` is that you have a license to the template code (by way of how the tool is licensed) and the template code was written by a human so licensable by them and then minimally modified by the tool.

Post reply on HN