Earlier quoted context omitted.
Yeah but did he die before anybody actually knew about it?
Is lib still around anymore. I can't find any functioning urls
Anthropic agrees to pay $1.5B to settle lawsuit with book authors
91–100 of 761 posts
Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#92I wonder who will be the first country to make an exception to copyright law for model training libraries to attract tax revenue like Ireland did for tech companies in the EU. Japan is part of the way there, but you couldn't do a common crawl type thing. You could even make it a library of congress type of setup.
As long as you're not distributing, it's legal in Switzerland to download copyrighted material. (Switzerland was on the naughty US/MPAA list for a while, might still be)
Or what if not even distributing it but rather distributing the outputs of the LLM (so closed source LLM like anthropic)
I am genuinely curious as to if there is some gray area that might be exploited by AI companies as I am pretty sure that they don't want to pay 1.5B dollars yet still want to exploit the works of authors. (let's call a spade a spade)
Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#93Settlement Terms (from the case pdf) 1. A Settlement Fund of at least $1.5 Billion: Anthropic has agreed to pay a minimum of $1.5 billion into a non-reversionary fund for the class members. With an estimated 500,000 copyrighted works in the class, this would amount to an approximate gross payment of $3,000 per work. If the final list of works exceeds 500,000, Anthropic will add $3,000 for each additional work. 2. Des…
So... it would be a lot cheaper to just buy all of the books?
A settlement means the claimants no longer have a claim, which means if they're also part of- say, the New York Times affiliated lawsuit- they have to withdraw. A neat way of kneecapping a country wide decision that LLM training on copy written material is subject to punitive measures don't you think?
Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#94One thing that comes to mind is... Is there a way to make your content on the web "licensed" in a way where it is only free for human consumption? I.e. effectively making the use of AI crawlers pirating, thus subject to the same kind of penalties here?
The purpose of the copyright protections is to promote "sciences and useful arts," and the public utility of allowing academia to investigate all works(1) exceeds the benefits of letting authors declare their works unponderable to the academic community.
(1) And yet, textbooks are copyrighted and the copyright is honored; I'm not sure why the academic fair-use exception doesn't allow scholars to just copy around textbooks without paying their authors.
Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#95Earlier quoted context omitted.
Paying $3,000 for pirating a ~$30 book seems disproportionate.
Not if 100 companies did it and they all got away. This is to teach a lesson because you cannot prosecute all thieves. Yale Law Journal actually writes about this, the goal is to deter crime because in most cases damages cannot be recovered or the criminal will never be caught in the first place.
Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#96To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.
Paying $3,000 for pirating a ~$30 book seems disproportionate.
It’s crazy to imagine, but there was surely a document or slack message thread discussing where to get thousands of books, and they just decided to pirate them and that was OK. This was entirely a decision based on ease or cost, not based on the assumption it was legal. Piracy can result in jail time IIRC, so honestly it’s lucky the employee who suggested this, or took the action avoided direct legal liability.
Oh and I’m pretty sure other companies (meta) are in litigation over this issue, and the publishers knew that settlement below the full legal limit would limit future revenue.
Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#97One thing that comes to mind is... Is there a way to make your content on the web "licensed" in a way where it is only free for human consumption? I.e. effectively making the use of AI crawlers pirating, thus subject to the same kind of penalties here?
Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#98Earlier quoted context omitted.
The blogger’s content was freely available, this fine is for piracy.
This is not a fine, it's a settlement to recompense authors. More broadly, I think that's a goofy argument. The books were "freely available" too. Just because something is out there, doesn't necessarily mean you can use it however you want, and that's the crux of the debate.
Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#99To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.
> It’s important in the fair use assessment to understand that the training itself is fair use, I think that this is a distinction many people miss. If you take all the works of Shakespeare, and reduce it to tokens and vectors is it Shakespeare or is it factual information about Shakespeare? It is the latter, and as much as organizations like the MLB might want to be able to copyright a fact you simply cannot do that…
This seems too cute by half, courts are generally far more common sense than that in applying the law.
This is like saying using `rails generate model:example` results in a bunch of code that isn't yours, because the tool generated it according to your specifications.
Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors
#100To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.
> It’s important in the fair use assessment to understand that the training itself is fair use, I think that this is a distinction many people miss. If you take all the works of Shakespeare, and reduce it to tokens and vectors is it Shakespeare or is it factual information about Shakespeare? It is the latter, and as much as organizations like the MLB might want to be able to copyright a fact you simply cannot do that…