A one time payment like 1.5B doesn’t do anything. There needs to be a royalty payment based on if the AI regurgitates existing ideas. That is probably the correct way to legislate this. If anything a human does can instantly be copied by an LLM, and then sent to all its subscribers, things need to change
Judge approves $1.5B Anthropic settlement for pirated books used to train Claude
491–500 of 654 posts
Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude
#492Earlier quoted context omitted.
Can you cite any information on this not being possible for the vast majority? Or is it simply that the correct prompt hasn't been written for all possible cases? I also fail to see the difference if logic/harnessing is added around a vector database that can output the complete corpus, but simply is instructed not to. It very clearly is still compressing the information into the vector weights, and then recovering t…
Information entropy. The amount of data an LLM ingests cannot be compressed to the size of the weights even at maximum theoretical compression.
LLMs are obviously capable of producing "exact" phrases as well. Ask it to give you famous quotes, it can do it. Ask it to read a paper for you and cite it, it can do it.
Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude
#493Earlier quoted context omitted.
> So will you owe life long compensation for all the knowledge you got from books too? You're just falling into the trap of anthropomorphizing the phrase "training" in the context of LLMs, which is not the same things as what humans do. There is no evidence they are the same thing and there is nothing to support the notion that what an LLM does when it "trains" on a book is equivalent to a human reading it.
Whatever "training" is, if you can't persuade the machine to spit substantially the same text back out verbatim, it's clearly not something that falls under copy right law either, because there's no copy. Yes, for some texts that's possible. But for the vast majority, it is not.
The New York Times lawsuit is resting on the point that large chunks of undigested articles can be vomited out. OpenAI tried to have the lawsuit thrown out but the courts permitted it to continue.
The Times... alleged that OpenAI's ChatGPT and Microsoft's Copilot had produced near-verbatim replicas of copyrighted articles, that the chatbots generated hallucinated content falsely attributed to the Times, ...
https://en.wikipedia.org/wiki/The_New_York_Times_v._Microsof...
Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude
#494Earlier quoted context omitted.
You’re arguing that freely remixing original work will give rise to greatness that’s even better than original work?
Isn't most creative work synthesis rather than unique whole-cloth creation? Look at what happens with software when it is open sourced and allowed to be remixed freely. Are we better or worse off because of it?
With open source, I should note, its remixing is in fact governed by copyright.
Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude
#495Earlier quoted context omitted.
> As far as I'm concerned, the courts are wrong, and training on ill gotten copyrighted material is not fair use. It’s important to remember that a court’s job is to apply law to a situation. When a court gets something wrong it’s a misinterpretation of the law and will, by definition, be overturnable on appeal. I suspect that your objection isn’t that the court is wrong, it’s that the law is wrong.
It's not a settled area of law and there is a SDNY judge that has a completely different application of the fair use analysis in the same exact context and came to a completely different conclusion (that it is not fair use).
Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude
#496Earlier quoted context omitted.
> but illegal to train on output of LLMs. Since when?
Typically the big LLM providers write in the their ToS that it is prohibited to use their output to train another LLM. Whereas for a book it is fair use.
Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude
#497A one time payment like 1.5B doesn’t do anything. There needs to be a royalty payment based on if the AI regurgitates existing ideas. That is probably the correct way to legislate this. If anything a human does can instantly be copied by an LLM, and then sent to all its subscribers, things need to change
I agree. If I pirate a book and share it on the web, and I get busted for doing so, and subsequently pay a fine, I don't get to KEEP sharing it on the web. Now, if I license the book, I might be able to come to an agreement with the author/publisher whereby I can share some of it.
(Then again, US copyright law often is insanely draconian.)
Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude
#498A one time payment like 1.5B doesn’t do anything. There needs to be a royalty payment based on if the AI regurgitates existing ideas. That is probably the correct way to legislate this. If anything a human does can instantly be copied by an LLM, and then sent to all its subscribers, things need to change
Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude
#499Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude
#500Earlier quoted context omitted.
Intellectual property protections have a finite lifespan to begin with, and as a tenet of Western Civilization are barely 300 years old. For works published before copyright laws existed or after property protections expire, anyone should be able to use it for anything forever without consent. Intellectual work still manages to get funded in this 'insane world' - although given the classical artist/patron system has…
Are you really saying that because IP protections are (and should be, I agree) time-limited, they're not doing anything in the first place? This is patently ridiculous. > Speaking of insane worlds, how does the concept of the Public Domain work in yours? Can you elaborate on what you're asking? I don't understand your question.
1. Copyright protections as a concept are an incredibly modern phenomenon, mostly limited in practice to Western Capitalist Democracies. 2. Outside of a short monetisable window (albeit one extended and irrevocably marred by Disney/Sonny Bono) your 'insane' hypothesis is in fact the status quo 3. Much intellectual work is published into the Public Domain, and all copyrighted work eventually ends up in the Public Domain. Your position appears to presuppose a world without such an entity.
As to what copyright actually achieves? It's mostly a mechanism by which the media gatekeepers and owners of capital use legislative and social imbalance of power to deny artist the rights and royalties for mechanical reproduction and otherwise impose financial serfdom.
This is achieved mainly by Copyright Enclosure, whereby musicians are typically pressured or contractually obligated to surrender their master recordings and intellectual property, and by contractual clauses like Controlled Composition Clauses, whereby Labels reduce the mechanical royalties they pay to artists who write their own songs, often paying below the standard statutory rate.