Live data from Hacker News

Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

apnews.com

491–500 of 654 posts

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#491

A one time payment like 1.5B doesn’t do anything. There needs to be a royalty payment based on if the AI regurgitates existing ideas. That is probably the correct way to legislate this. If anything a human does can instantly be copied by an LLM, and then sent to all its subscribers, things need to change

I'd rather see AI studios do the same as the film industry, pay a one-time up front cost per major model (or major.minor?) depending on how they contract it out. This also allows smaller startups to license books for less than a larger frontier studio would. In theory and hopefully, the pricing would not be too insane per book, you want them to rent more books and spend more, not go back to pirating right?

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#492

Earlier quoted context omitted.

Can you cite any information on this not being possible for the vast majority? Or is it simply that the correct prompt hasn't been written for all possible cases? I also fail to see the difference if logic/harnessing is added around a vector database that can output the complete corpus, but simply is instructed not to. It very clearly is still compressing the information into the vector weights, and then recovering t…

Information entropy. The amount of data an LLM ingests cannot be compressed to the size of the weights even at maximum theoretical compression.

Is that relevant? I can use a lossy compression algorithm such that the original could never be recovered from the image I've produced, but that derived image would surely be under copyright.

LLMs are obviously capable of producing "exact" phrases as well. Ask it to give you famous quotes, it can do it. Ask it to read a paper for you and cite it, it can do it.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#493
post #453

Earlier quoted context omitted.

> So will you owe life long compensation for all the knowledge you got from books too? You're just falling into the trap of anthropomorphizing the phrase "training" in the context of LLMs, which is not the same things as what humans do. There is no evidence they are the same thing and there is nothing to support the notion that what an LLM does when it "trains" on a book is equivalent to a human reading it.

Whatever "training" is, if you can't persuade the machine to spit substantially the same text back out verbatim, it's clearly not something that falls under copy right law either, because there's no copy. Yes, for some texts that's possible. But for the vast majority, it is not.

> spitting out verbatim text

The New York Times lawsuit is resting on the point that large chunks of undigested articles can be vomited out. OpenAI tried to have the lawsuit thrown out but the courts permitted it to continue.

The Times... alleged that OpenAI's ChatGPT and Microsoft's Copilot had produced near-verbatim replicas of copyrighted articles, that the chatbots generated hallucinated content falsely attributed to the Times, ...

https://en.wikipedia.org/wiki/The_New_York_Times_v._Microsof...

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#494

Earlier quoted context omitted.

You’re arguing that freely remixing original work will give rise to greatness that’s even better than original work?

Isn't most creative work synthesis rather than unique whole-cloth creation? Look at what happens with software when it is open sourced and allowed to be remixed freely. Are we better or worse off because of it?

There's nothing that prevents people from remixing things that are not copyrighted and create something amazing that others are interested in or of cultural value.

With open source, I should note, its remixing is in fact governed by copyright.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#495

Earlier quoted context omitted.

> As far as I'm concerned, the courts are wrong, and training on ill gotten copyrighted material is not fair use. It’s important to remember that a court’s job is to apply law to a situation. When a court gets something wrong it’s a misinterpretation of the law and will, by definition, be overturnable on appeal. I suspect that your objection isn’t that the court is wrong, it’s that the law is wrong.

It's not a settled area of law and there is a SDNY judge that has a completely different application of the fair use analysis in the same exact context and came to a completely different conclusion (that it is not fair use).

I would like to see a citation on that b/c I am unaware of it. The only case I see in SDNY is the NYT v OpenAI case which has not been ruled on yet. https://www.reuters.com/legal/legalindustry/copyright-law-20...

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#496
post #377

Earlier quoted context omitted.

> but illegal to train on output of LLMs. Since when?

Typically the big LLM providers write in the their ToS that it is prohibited to use their output to train another LLM. Whereas for a book it is fair use.

ToS are usually not worth the toilet paper they're printed on. They're not legally binding.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#497

A one time payment like 1.5B doesn’t do anything. There needs to be a royalty payment based on if the AI regurgitates existing ideas. That is probably the correct way to legislate this. If anything a human does can instantly be copied by an LLM, and then sent to all its subscribers, things need to change

I agree. If I pirate a book and share it on the web, and I get busted for doing so, and subsequently pay a fine, I don't get to KEEP sharing it on the web. Now, if I license the book, I might be able to come to an agreement with the author/publisher whereby I can share some of it.

The post specifically proposes royalties for ideas from books, not royalties for the books themselves. You would absolutely still be able to share ideas you learned from the books you pirated in that situation. It'd be insanely draconian if you couldn't.

(Then again, US copyright law often is insanely draconian.)

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#498

A one time payment like 1.5B doesn’t do anything. There needs to be a royalty payment based on if the AI regurgitates existing ideas. That is probably the correct way to legislate this. If anything a human does can instantly be copied by an LLM, and then sent to all its subscribers, things need to change

how would you do that? you can copyright words, but you can't copyright an idea (you can patent some ideas, but not all of them)

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#500
post #460

Earlier quoted context omitted.

Intellectual property protections have a finite lifespan to begin with, and as a tenet of Western Civilization are barely 300 years old. For works published before copyright laws existed or after property protections expire, anyone should be able to use it for anything forever without consent. Intellectual work still manages to get funded in this 'insane world' - although given the classical artist/patron system has…

Are you really saying that because IP protections are (and should be, I agree) time-limited, they're not doing anything in the first place? This is patently ridiculous. > Speaking of insane worlds, how does the concept of the Public Domain work in yours? Can you elaborate on what you're asking? I don't understand your question.

Your contention was that it was an insane position that anyone should be able to use a piece of literature or a song for anything forever without consent. I simply highlighted the absurdity of that based on the fact that:

1. Copyright protections as a concept are an incredibly modern phenomenon, mostly limited in practice to Western Capitalist Democracies. 2. Outside of a short monetisable window (albeit one extended and irrevocably marred by Disney/Sonny Bono) your 'insane' hypothesis is in fact the status quo 3. Much intellectual work is published into the Public Domain, and all copyrighted work eventually ends up in the Public Domain. Your position appears to presuppose a world without such an entity.

As to what copyright actually achieves? It's mostly a mechanism by which the media gatekeepers and owners of capital use legislative and social imbalance of power to deny artist the rights and royalties for mechanical reproduction and otherwise impose financial serfdom.

This is achieved mainly by Copyright Enclosure, whereby musicians are typically pressured or contractually obligated to surrender their master recordings and intellectual property, and by contractual clauses like Controlled Composition Clauses, whereby Labels reduce the mechanical royalties they pay to artists who write their own songs, often paying below the standard statutory rate.

Post reply on HN