Live data from Hacker News

Anthropic agrees to pay $1.5B to settle lawsuit with book authors

nytimes.com

31–40 of 761 posts

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#31
I'm gonna say one thing. If you agree that something was unfairly taken from book authors, then the same thing was taken from people publishing on the web, and on a larger scale.

Book authors may see some settlement checks down the line. So might newspapers and other parties that can organize and throw enough $$$ at the problem. But I'll eat my hat if your average blogger ever sees a single cent.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#32
For legal observers, Judge William Haskell Alsup’s razor-sharp distinction between usage and acquisition is a landmark precedent: it secures fair use for transformative generative AI while preserving compensation for copyright holders. In a just world, this balance would elevate him to the highest court of the land, but we are far from a just world.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#33
post #8

To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.

Google scanned many books quite a while ago, probably way more than LibGen. Are they good to use them for training?

I imagine the problem there is they primarily scanned library books so I doubt they have the same copyright protections here as if they bought them

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#36
post #17
post #8

To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.

Wdym Rainbows End was prescient?

There's a scene early on where libraries are being destructively shredded, with the shreds scanned and reconstructed as digital versions.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#37
post #20

This settlement I guess could be a landmark moment. $1.5 billion is a staggering figure and I hope it sends a clear signal that AI companies can’t just treat creative work as free training data.

I mean the ruling does in fact find that treating this particular kind of creative work qualifies as fair use.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#38

They also agreed to destroy the pirated books. I wonder how large of a portion of their training data comes from these shadow libraries, and if AI labs in countries that have made it clear they won't enforce anti-piracy laws against AI companies will get a substantial advantage by continuing to use shadow libraries.

Perhaps they'll quickly rent the whole contents of a few physical libraries and then scan them all

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#40
post #8

To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.

Google scanned many books quite a while ago, probably way more than LibGen. Are they good to use them for training?

If they legally purchased them I dont think why not. IIRC they did borrow from libraries so probably not every book in Google Books
Post reply on HN