Live data from Hacker News

Anthropic agrees to pay $1.5B to settle lawsuit with book authors

nytimes.com

401–410 of 761 posts

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#401

Earlier quoted context omitted.

> which in my opionion is exactly what is wrong with practices like these. What's actually wrong with this? They paid $1.5B for a bunch of pirated books. Seems like a fair price to me, but what do I know. The settlement should reflect society's belief of the cost or deterrent, I'm not sure which (maybe both). This might be controversial, but I think a free society needs to let people break the rules if they are willi…

> What's actually wrong with this? It's because they did not choose to pay for the books; they were forced to pay and they would not have done so if the lawsuit had not fallen this way. If you are not sure why this is different from "they paid for pirated books (as if it were a transaction)", then this may reflect a lack of awareness of how fair exchange and trust both function in a society.

Settling is not forced

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#402
post #359
post #44

I wonder who will be the first country to make an exception to copyright law for model training libraries to attract tax revenue like Ireland did for tech companies in the EU. Japan is part of the way there, but you couldn't do a common crawl type thing. You could even make it a library of congress type of setup.

This is already a thing in several places. EU has copyright exemptions for AI training. You don't need to respect opt outs if you are doing research. South Korea, Japan has some exemptions too I think? Singapore has very strong copyright exemptions for AI training. You can completely ignore opt-outs legally, even if doing it commercially. Just search up "TDM laws globally".

So could they have library genesis on a local server and other pirate sources and use that for training data then? That is the level I'm speaking of, much like common crawl and the reddit archive

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#403
post #369
post #277

Earlier quoted context omitted.

> But nobody was ever going to that Didn't Google have a long standing project to do just that? https://en.wikipedia.org/wiki/Google_Books

Crazy to think we've been helping train AI through captchas long before the "click all squares containing" ones.

"stop spam. read books." is a very ironic phrase to look back on considering the amount of spam on the internet that LLMs have enabled

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#404

Earlier quoted context omitted.

> It was faster to just put unlicensed taxis on the streets and use investor money to pay fines and lobby for favorable legislation And thank god they did. There was no perfectly legal channel to fix the taxi cartel. Now you don't even have to use Uber in many of these places because taxis had to compete - they otherwise never would have stopped pulling the "credit card reader is broken" scam, taking long routes on p…

i dont know that its such a great thing in the end. Uber/Lyft is 50-100% more expensive now than taxis were before. Theyre entrenched in different ways.

How much of this is inflation?

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#405
post #327

Earlier quoted context omitted.

If I'm reading this right yes the training was fair use, but I was responding (unclearly) to the claim that the pirated books weren't used to train commercially released LLMs. The judge complained that it wasn't clear what was actually used, from the June order https://fingfx.thomsonreuters.com/gfx/legaldocs/jnvwbgqlzpw/... [pdf]: > Notably, in its motion, Anthropic argues that pirating initial copies of Authors’ boo…

Thanks for this info. I was looking for which pirated books were used for which model. Ethically speaking, if Anthropic (a) did later purchase every book it pirated or (b) compensated every author whose book was pirated, would it absolve an illegally trained model of its "sins"? To me, the taint still remains. Which is a shame, because it's considered the best coding model so far.

> Ethically speaking, if Anthropic (a) did later purchase every book it pirated or (b) compensated every author whose book was pirated, would it absolve an illegally trained model of its "sins"?

No, it part because it removes agency from the authors/rightsholders. Maybe they don't want to sell Anthropic their books, maybe they want royalties, etc.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#406
post #8

To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.

[dead]

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#407
post #299
post #248

I can't help but feel like this is a huge win for Chinese AI. Western companies are going to be limited in the amount of data they can collect and train on, and Chinese (or any foreign AI) is going to have access to much more and much better data.

The West can end the endless pain and legal hurdles to innovation by limiting the copyright. They can do it if there is will to open up the gates of information to everyone. The duration of 70 years after death of the author or 90 years for companies is excessively long. It should be ~25 years. For software it should be 10 years. And if AI companies want recent stuff, they need to pay the owners. However, the West wa…

The vast majority of books don't generate any profits past the first few years, so I prefer Lawrence Lessig's proposal of copyright renewal at five-year intervals with a fee. Under this scheme, most books would enter the public domain after five years

https://www.econlib.org/library/Columns/y2003/Lessigcopyrigh...

Lessig: Not for this length of time, no. Copyright shouldn’t be anywhere close to what it is right now. In my book I proposed a system where you’d have to renew after every five years and you get a maximum term of 75 years. I thought that was pretty radical at the time. The Economist, after the Eldred decision, came out with a proposal—let’s go back to 14 years, renewable to 28 years. Nobody needs more than 14 years to earn the return back from whatever they produced.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#408
post #320

Earlier quoted context omitted.

By US law, cccording to Author's Guild vs Google[1] on the Google book scanning project, scanning books for indexes is fair use. Additionally: > Every human has the right to read those books. Since when? I strongly disagree - knowledge should be free. I don't think the author's arrangement of the words should be free to reproduce (ie, I think some degree of copyright protection is ethical) but if I want to use a tool…

Knowledge should be free. Unfortunately, OpenAI and most other AI companies are for-profit, and so they vacuum up the commons, and produce tooling which is for-profit. If you use the commons to create your model, perhaps you should be obligated to distribute the model for free (or I guess for the cost of distribution) too.

> vacuum up the commons

A vacuum removes what it sucks in. The commons are still as available as they ever were, and the AI gives one more avenue of access.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#409
post #228

Earlier quoted context omitted.

> Buying used copies of books, scanning them, and training on it is fine. But nobody was ever going to that, not when there are billions in VC dollars at stake for whoever moves fastest. Everybody will simply risk the fine, which tends to not be anywhere close to enough to have a deterrent effect in the future. That is like saying Uber would have not had any problems if they just entered into a licensing contract wit…

> It was faster to just put unlicensed taxis on the streets and use investor money to pay fines and lobby for favorable legislation And thank god they did. There was no perfectly legal channel to fix the taxi cartel. Now you don't even have to use Uber in many of these places because taxis had to compete - they otherwise never would have stopped pulling the "credit card reader is broken" scam, taking long routes on p…

[deleted]

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#410
It is a good opportunity to ask: is it true, that Anthropic can get indemnification from user actions that end up in the company being sued? User actions that are related to the use of Claude. Even just for the accusation. The user needs to cover their bills of lawyers and proceedings. Also they take control of the legal process, can do the way they please, settle or what, user footing the bill. Without limit. Be the user an individual or organization, doesn't matter.

Sounds harsh, if true. Making its use practical only for hobby projects basically where the results of Claude kept for yourself completely (be it information, product using Claude, or product is made by using Claude). Difficult to believe, I hope I heard it wrong.

Post reply on HN