Earlier quoted context omitted.
> which in my opionion is exactly what is wrong with practices like these. What's actually wrong with this? They paid $1.5B for a bunch of pirated books. Seems like a fair price to me, but what do I know. The settlement should reflect society's belief of the cost or deterrent, I'm not sure which (maybe both). This might be controversial, but I think a free society needs to let people break the rules if they are willi…
> I think a free society needs to let people break the rules if they are willing to pay the cost so you don't think super rich people should be bound by laws at all? Unless you made the cost proportional to (maybe expontial to) somebody's wealth, you would be creating a completely lawless class who would wreak havoc on society.
Anthropic agrees to pay $1.5B to settle lawsuit with book authors
461–470 of 761 posts
Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#462This is sad for open source AI, piracy for the purpose of model training should also be fair use because otherwise only the big companies who can afford to pay off publishers like Anthropic will be able to do so. There is no way to buy billions of books just for model training, it simply can't happen.
Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#463Settlement Terms (from the case pdf) 1. A Settlement Fund of at least $1.5 Billion: Anthropic has agreed to pay a minimum of $1.5 billion into a non-reversionary fund for the class members. With an estimated 500,000 copyrighted works in the class, this would amount to an approximate gross payment of $3,000 per work. If the final list of works exceeds 500,000, Anthropic will add $3,000 for each additional work. 2. Des…
I’m an author, can I get in on this?
Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#464Earlier quoted context omitted.
> But nobody was ever going to that Didn't Google have a long standing project to do just that? https://en.wikipedia.org/wiki/Google_Books
This lawsuit also makes sure that only parties that can train an AI with good enough training material are now - Google - Anthropic - Any Chinese company who do not care about copyright laws What is the cost of buying and scanning books? Copyright law needs to be fixed and its ridiculous hundred years tenure chopped away.
Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#465To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.
> It’s important in the fair use assessment to understand that the training itself is fair use Is this completely settled legally? It is not obvious to me it would be so
Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#466Earlier quoted context omitted.
> Buying used copies of books, scanning them, and training on it is fine. But nobody was ever going to that, not when there are billions in VC dollars at stake for whoever moves fastest. Everybody will simply risk the fine, which tends to not be anywhere close to enough to have a deterrent effect in the future. That is like saying Uber would have not had any problems if they just entered into a licensing contract wit…
> But nobody was ever going to that Didn't Google have a long standing project to do just that? https://en.wikipedia.org/wiki/Google_Books
The Google Books project also faced a copyright lawsuit, which was eventually decided in favor of Google.
After contacting major publishers about possibly licensing their books, [former head of the Google Books project] bought physical books in bulk from distributors and retailers, according to court documents. He then hired outside organizations to dissemble the books, scan them and create digital copies that could be used to train the company’s AI. technologies.
Judge Alsup ruled that this approach was fair use under the law. But he also found the company’s previous approach — downloading and storing books from shadow libraries like Library Genesis and Pirate Library Mirror — was illegal.Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#467To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.
I think the jury is still out on how fair use applies to AI. Fair use was not designed for what we have now. I could read a book, but its highly unlikely I could regurgitate it, much less months or years later. An LLM, however, can. While we can say "training is like reading", its also not like reading at all due to permanent perfect recall. Not only does an LLM have perfect recall, it also has the ability to distrib…
The way this technology is being used clearly violates the intent behind copyright law, it undermines its goals and results in harm that it was designed to prevent. I believe that doing this without extensive public discussion and consensus is anti-democratic.
We always end up discussing concrete implementation details of how copyright is currently enforced, never the concept itself. Is there a good word for this? Reification?
Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#468From a systems design perspective, $3,000 per book makes this approach completely unscalable compared to web scraping. It's like choosing between a O(n) and O(n²) algorithm - legally compliant data acquisition has fundamentally different scaling characteristics than the 'move fast and break things' approach most labs took initially.
I don't know if anyone has actually read the article or the ruling, but this is about pirating books. Anthropic went back and bought->scanned->destroyed physical copies of them afterward... but they pirated them first, and that's what this settlement is about. The judge also said: > “The training use was a fair use,” he wrote. “The technology at issue was among the most transformative many of us will see in our lifet…
Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#469Earlier quoted context omitted.
> But nobody was ever going to that Didn't Google have a long standing project to do just that? https://en.wikipedia.org/wiki/Google_Books
This lawsuit also makes sure that only parties that can train an AI with good enough training material are now - Google - Anthropic - Any Chinese company who do not care about copyright laws What is the cost of buying and scanning books? Copyright law needs to be fixed and its ridiculous hundred years tenure chopped away.
> Anthropic also agreed to delete the pirated works it downloaded and stored.
Also > As part of the settlement, Anthropic said that it did not use any pirated works to build A.I. technologies that were publicly released.Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#470Earlier quoted context omitted.
i dont know that its such a great thing in the end. Uber/Lyft is 50-100% more expensive now than taxis were before. Theyre entrenched in different ways.
Idk how it is in the US but in eastern Europe that's only true if surge is on and even so considering how shitty the quality of service was before Uber it's fine.