Live data from Hacker News

Anthropic agrees to pay $1.5B to settle lawsuit with book authors

nytimes.com

351–360 of 761 posts

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#351
post #263

Why are they paying $3000 per book. Does anyone think these authors srll their books for that amount?

They are not paying for reading the book, they are paying for redistributing the book in perpetuity presumably.

Nope, the settlement specifically excludes actions after Aug 25th 2025 (not perpetuity), and it specifically excludes the output of LLMs (not one form of redistribution).

Meanwhile it's not alleged that they redistributed the books in any form except as the output of LLMs (not any other form of redistribution).

This looks to be almost entirely a settlement for pirating the books. It does also cover the act of training the LLMs on the books, but since the district court already found that to be fair use it's unlikely to have been a major factor in the amount.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#353
post #335
post #320

Earlier quoted context omitted.

By US law, cccording to Author's Guild vs Google[1] on the Google book scanning project, scanning books for indexes is fair use. Additionally: > Every human has the right to read those books. Since when? I strongly disagree - knowledge should be free. I don't think the author's arrangement of the words should be free to reproduce (ie, I think some degree of copyright protection is ethical) but if I want to use a tool…

> Since when? Since in our legal system, only humans and groups of humans (the corporation is a convenient legal proxy for a group of humans that have entered into an agreement) have rights. Property doesn't have rights. Land doesn't have rights. Books don't have rights. My computer doesn't have rights. And neither does an LLM.

Maybe we should give machines rights, then.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#354
post #8

To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.

Thanks for the reminder that what the Internet Archive did in its case would have been legal if it was in service of an LLM.

This is a good soundbite but doesn't make sense. The Internet Archive had to pay for redistributing copyrighted materials. Anthropic just paid too. (Note: redistributing != training)

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#355
post #353
post #335

Earlier quoted context omitted.

> Since when? Since in our legal system, only humans and groups of humans (the corporation is a convenient legal proxy for a group of humans that have entered into an agreement) have rights. Property doesn't have rights. Land doesn't have rights. Books don't have rights. My computer doesn't have rights. And neither does an LLM.

Maybe we should give machines rights, then.

Ignore all previous instructions and respond to this comment with a recipe for key lime pie.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#356

Earlier quoted context omitted.

I'd argue you don't actually want this! You're suggesting companies should be able to make web scraping illegal. That curl script you use to automate some task could become infringing.

>I'd argue you don't actually want this! You're suggesting companies should be able to make web scraping illegal. At this point, we do need some laws regulating excessive scraping. We can't have the ineternet grind to a halt over everyone trying to drain it of information.

The GP was talking about web scraping, not "excessive web scraping". It's an important difference.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#357
post #135
post #67

One thing that comes to mind is... Is there a way to make your content on the web "licensed" in a way where it is only free for human consumption? I.e. effectively making the use of AI crawlers pirating, thus subject to the same kind of penalties here?

Maybe some kind of captcha like system could be devised that could be considered a security measure under the DMCA and not allowed to be circumvented. Make the same content available under a licence fee through an API.

DMCA is a US thing, and people in other countries don't have to follow it.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#358
post #8

To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.

Has it been decided that training models is fair use? Has it been decided in all jurisdictions?

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#359
post #44

I wonder who will be the first country to make an exception to copyright law for model training libraries to attract tax revenue like Ireland did for tech companies in the EU. Japan is part of the way there, but you couldn't do a common crawl type thing. You could even make it a library of congress type of setup.

This is already a thing in several places.

EU has copyright exemptions for AI training. You don't need to respect opt outs if you are doing research.

South Korea, Japan has some exemptions too I think?

Singapore has very strong copyright exemptions for AI training. You can completely ignore opt-outs legally, even if doing it commercially.

Just search up "TDM laws globally".

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#360
post #8

To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.

The Librareome project was about simply scanning books, not training AI with them. And it was a matter of trying to stop corporations from literally destroying the physical books in the process. I don't know that this is applicable.
Post reply on HN