Live data from Hacker News

Anthropic agrees to pay $1.5B to settle lawsuit with book authors

nytimes.com

221–230 of 761 posts

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#221
post #25

Earlier quoted context omitted.

> Rainbows End was prescient in many ways. Agreed. Great book for those looking for a read: https://www.goodreads.com/book/show/102439.Rainbows_End The author, Vernor Vinge, is also responsible for popularizing the term 'singularity'.

RIP to the legend. He has a lot of really fun ideas spread across his books.

I didn't realize Vernor Vinge had passed away... Sad TIL

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#222
post #140

Earlier quoted context omitted.

Printing press, audio recording, movies, radio, television were also transformative. Did not get rid of copyright or actually brought them. I feel it is insane that authors do not receive some sort of standard compensation for each training use. Say a few hundred to a few thousand depending on complexity of their work.

Why would they earn more from models reading their works than I would pay to read it?

Same reason why the enterprise edition is more expensive than personal. Companies have more money to give and usually use it to generate profit. Individuals do not.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#223

Settlement Terms (from the case pdf) 1. A Settlement Fund of at least $1.5 Billion: Anthropic has agreed to pay a minimum of $1.5 billion into a non-reversionary fund for the class members. With an estimated 500,000 copyrighted works in the class, this would amount to an approximate gross payment of $3,000 per work. If the final list of works exceeds 500,000, Anthropic will add $3,000 for each additional work. 2. Des…

Only 500,000 copyrighted works?

I was under the impression they had downloaded millions of books.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#224
post #174

Earlier quoted context omitted.

Because ones doing the training are profiting from it. Ai is not a human with limited time. And it is also owned by a company not a legal person. I might find argument of comparing it to human when it is fully legal person and cutting power to it or deleting is treated as murder. Before that it is just bullshit. And fundamentally reason for copy right to exist is to support creators and to promote them to create more…

If I buy a book, learn something, and then profit from it, should I also be paying more than the original price to read the book? > Ai is not a human with limited time AI is also bound by time, physics, and limited capacity. It does certain things better or faster than us, it fails miserably at certain things we don't even think about being complex (like opening a door) > And it is also owned by a company not a legal…

> If I buy a book, learn something, and then profit from it, should I also be paying more than the original price to read the book?

Depends on how you do it. Clearly reading the book word from word is different from making a podcast talking about your interpretation of the book.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#225

Earlier quoted context omitted.

RIP to the legend. He has a lot of really fun ideas spread across his books.

I didn't realize Vernor Vinge had passed away... Sad TIL

There was a nice discussion & nostalgia at the time (1151 points, 2024, 320 comments) https://news.ycombinator.com/item?id=39775304

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#226

Earlier quoted context omitted.

Scanning (copying) is¹ not allowed. Reading is. What is in a library, you can freely read. Find the most appropriate way. You do not need to have bought the book. ¹(Edit: or /may/ not be allowed, see posts below.)

There are no terms and conditions attached to library books beyond copyright law (which says nothing about scanning) and the general premise of being a library (return the book in good condition on time or pay).

Copyright law in the USA may be more liberal about scanning than other jurisdictions (see the parallel comment from gpm), which expressly regulate the amount of copying of material you do not own as an item.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#228
post #8

To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.

> Buying used copies of books, scanning them, and training on it is fine.

But nobody was ever going to that, not when there are billions in VC dollars at stake for whoever moves fastest. Everybody will simply risk the fine, which tends to not be anywhere close to enough to have a deterrent effect in the future.

That is like saying Uber would have not had any problems if they just entered into a licensing contract with taxi medallion holders. It was faster to just put unlicensed taxis on the streets and use investor money to pay fines and lobby for favorable legislation. In the same way, it was faster for Anthropic to load up their models with un-DRM'd PDFs and ePUBs from wherever instead of licensing them publisher by publisher.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#230
post #8

To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.

This is excellent news because it means that folks who pay for printed books and scan them also can train with their content. It's been said already that we've already trained on "the entire (public) internet." Printed books still hold a wealth of knowledge that could be useful in training models. And cheap, otherwise unwanted copies make great fodder for "destructive" scanning where you cut the spine off and feed it to a page scanner. There are online services that offer just that.
Post reply on HN