Anyone have a link to the class action? I published a book and would love to know if I'm in the class.
Anthropic agrees to pay $1.5B to settle lawsuit with book authors
191–200 of 761 posts
Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#192Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#193Earlier quoted context omitted.
Buying used copies of books, scanning them, training an employee with the scans: fair use. Unless legislation changes, model training is pretty much analogous to that. Now of course if the employee in question - or the LLM - regurgitates a copyrighted piece verbatim, that is a violation and would be treated accordingly in either case.
> Buying used copies of books, scanning them, training an employee with the scans: fair use. Does this still hold true if multiple employees are "trained" from scanned copies at the same time?
Regardless, the issue could be resolved by buying as many copies as you have concurrent model training instances. It isn't really an issue with training on copyrighted work, just a matter of how you do so.
Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#194Earlier quoted context omitted.
> It’s important in the fair use assessment to understand that the training itself is fair use, I think that this is a distinction many people miss. If you take all the works of Shakespeare, and reduce it to tokens and vectors is it Shakespeare or is it factual information about Shakespeare? It is the latter, and as much as organizations like the MLB might want to be able to copyright a fact you simply cannot do that…
> And what about all the other stuff that LLM's spit out? Who owns that. Well at present, no one. If you train a monkey or an elephant to paint, you cant copyright that work because they aren't human, and neither is an LLM. This seems too cute by half, courts are generally far more common sense than that in applying the law. This is like saying using `rails generate model:example` results in a bunch of code that isn'…
Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#195To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.
Paying $3,000 for pirating a ~$30 book seems disproportionate.
If anything it's too little based on precedent.
Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#196To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.
Thanks for the reminder that what the Internet Archive did in its case would have been legal if it was in service of an LLM.
Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#197Wait so they raised all that money just to give it to publishers? Can only imagine the pitch, yes please give us billions of dollars. We are going to make a huge investment like paying of our lawsuits.
So long as there is an excuse to justify money flows, that's fine, big capital doesn't really care about the excuse; so long as the excuse is just persuasive enough to satisfy the regulators and the judges.
Money flows happen independently, then later, people try to come up with good narratives. This is exactly what happened in this case. They paid the authors a lot of money as a settlement and agreed on a narrative which works for both sets of people; that training was fine, it's the pirating which was a problem...
It's likely why they settled; they preferred to pay a lot of money and agree on some false narrative which works for both groups rather than setting a precedent that AI training on copyrighted material is illegal; that would be the biggest loss for them.
Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#198Earlier quoted context omitted.
> It’s important in the fair use assessment to understand that the training itself is fair use, I think that this is a distinction many people miss. If you take all the works of Shakespeare, and reduce it to tokens and vectors is it Shakespeare or is it factual information about Shakespeare? It is the latter, and as much as organizations like the MLB might want to be able to copyright a fact you simply cannot do that…
Yes. Someone on this post mentioned that switzerland allows downloading copyrightable material but not distributing them. So things get even more dark because what becomes distribution can have a really vague definition and maybe the AI companies will only follow the law just barely, just for the sake of not getting hit with a lawsuit like this again. But I wonder if all this case did was maybe compensate the authors…
Note that the law specifically regulates software differently, so what you cannot do is just willy nilly pirate games and software.
What distribution means in this case is defined in the swiss law. However swiss law as a whole is in some ways vague, to leave a lot up to interpretation by the judiciary.
Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#199Earlier quoted context omitted.
Is it distribution though if someone trains a model in switzerland through downloading copyrighted material, training AI on it and then distributing it... Or what if not even distributing it but rather distributing the outputs of the LLM (so closed source LLM like anthropic) I am genuinely curious as to if there is some gray area that might be exploited by AI companies as I am pretty sure that they don't want to pay…
Using copyrighted material to train AI is a legal grey zone. The nyt vs openAI case is litigating this. The anthropic settlement here is about how the material is obtained. If openAI wins their case and switzerland rules the same way I dont think there would be a problem
We really are getting at some metaphysical / philosophical questions and maybe we will one day arrive at a question that just can't be answered (I think this is pretty close, right?) and then AI companies would do things freely without being accountable since sure you could take to the courts but how would you come to the decision...?
Another question though
So lets say that the nyt vs openAI case is going on, so in the meantime while they are litigating (lets say), could OpenAI still continue doing the same thing while the case is going on?
Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#200Earlier quoted context omitted.
Because ones doing the training are profiting from it. Ai is not a human with limited time. And it is also owned by a company not a legal person. I might find argument of comparing it to human when it is fully legal person and cutting power to it or deleting is treated as murder. Before that it is just bullshit. And fundamentally reason for copy right to exist is to support creators and to promote them to create more…
If I buy a book, learn something, and then profit from it, should I also be paying more than the original price to read the book? > Ai is not a human with limited time AI is also bound by time, physics, and limited capacity. It does certain things better or faster than us, it fails miserably at certain things we don't even think about being complex (like opening a door) > And it is also owned by a company not a legal…
Are they paying reasonable compensation? Say like with streaming services, movie theatres, radio and tv stations. As a whole their model is much close to those than individuals buying books, cds or dvds...
You might even consider Theatrical License or Public Performance License. Paid even if you have memorized a thing...
LLMs are just bad technology that require massive amount of inputs so the authors cannot be compensated enough for it. And I fully believe they should be. And lot more than single copy of their work under entirely ill-fitting first-sale doctrine does.