Live data from Hacker News

Anthropic agrees to pay $1.5B to settle lawsuit with book authors

nytimes.com

91–100 of 761 posts

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#91

Earlier quoted context omitted.

Yeah but did he die before anybody actually knew about it?

Is lib still around anymore. I can't find any functioning urls

There are mirrors on its' wikipedia page: https://en.wikipedia.org/wiki/Library_Genesis

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#92
post #69
post #44

I wonder who will be the first country to make an exception to copyright law for model training libraries to attract tax revenue like Ireland did for tech companies in the EU. Japan is part of the way there, but you couldn't do a common crawl type thing. You could even make it a library of congress type of setup.

As long as you're not distributing, it's legal in Switzerland to download copyrighted material. (Switzerland was on the naughty US/MPAA list for a while, might still be)

Is it distribution though if someone trains a model in switzerland through downloading copyrighted material, training AI on it and then distributing it...

Or what if not even distributing it but rather distributing the outputs of the LLM (so closed source LLM like anthropic)

I am genuinely curious as to if there is some gray area that might be exploited by AI companies as I am pretty sure that they don't want to pay 1.5B dollars yet still want to exploit the works of authors. (let's call a spade a spade)

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#93
post #77

Settlement Terms (from the case pdf) 1. A Settlement Fund of at least $1.5 Billion: Anthropic has agreed to pay a minimum of $1.5 billion into a non-reversionary fund for the class members. With an estimated 500,000 copyrighted works in the class, this would amount to an approximate gross payment of $3,000 per work. If the final list of works exceeds 500,000, Anthropic will add $3,000 for each additional work. 2. Des…

So... it would be a lot cheaper to just buy all of the books?

Few. This settlement potentially weakens all challenges to the use of copyrighted works in training LLM's. I'd be shocked if behind closed doors there wasn't some give and take on the matter between Executives/investors.

A settlement means the claimants no longer have a claim, which means if they're also part of- say, the New York Times affiliated lawsuit- they have to withdraw. A neat way of kneecapping a country wide decision that LLM training on copy written material is subject to punitive measures don't you think?

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#94
post #67

One thing that comes to mind is... Is there a way to make your content on the web "licensed" in a way where it is only free for human consumption? I.e. effectively making the use of AI crawlers pirating, thus subject to the same kind of penalties here?

I'm sure one can try, but copyright has all kinds of oddities and carve-outs that make this complicated. IANAL, but I'm fairly certain that, for example, if you tried putting in your content license "Free for all uses public and private, except academia, screw that ivory tower..." that's a sentiment you can express but universities are under no obligation legally to respect your wish to not have your work included in a course presentation on "wild things people put in licenses." Similarly, since the court has found that training an LLM on works is transformative, a license that says "You may use this for other things but not to train an LLM" couldn't be any more enforceable than a musician saying "You may listen to my work as a whole unit but God help you if I find out you sampled it into any of that awful 'rap music' I keep hearing about..."

The purpose of the copyright protections is to promote "sciences and useful arts," and the public utility of allowing academia to investigate all works(1) exceeds the benefits of letting authors declare their works unponderable to the academic community.

(1) And yet, textbooks are copyrighted and the copyright is honored; I'm not sure why the academic fair-use exception doesn't allow scholars to just copy around textbooks without paying their authors.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#95
post #89
post #68

Earlier quoted context omitted.

Paying $3,000 for pirating a ~$30 book seems disproportionate.

Not if 100 companies did it and they all got away. This is to teach a lesson because you cannot prosecute all thieves. Yale Law Journal actually writes about this, the goal is to deter crime because in most cases damages cannot be recovered or the criminal will never be caught in the first place.

If in most cases damages cannot be recovered or the criminal will never be caught in the first place, then what is the lesson being taught? Doesn't that just create a moral hazard where you "randomly" choose who to penalize?

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#96
post #68
post #8

To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.

Paying $3,000 for pirating a ~$30 book seems disproportionate.

I feel like proportionality is related also to the scale. If a student pirates a textbook, I’d agree that 100x is excessive, but this is a corporation handsomely profiting off of mass piracy.

It’s crazy to imagine, but there was surely a document or slack message thread discussing where to get thousands of books, and they just decided to pirate them and that was OK. This was entirely a decision based on ease or cost, not based on the assumption it was legal. Piracy can result in jail time IIRC, so honestly it’s lucky the employee who suggested this, or took the action avoided direct legal liability.

Oh and I’m pretty sure other companies (meta) are in litigation over this issue, and the publishers knew that settlement below the full legal limit would limit future revenue.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#97
post #67

One thing that comes to mind is... Is there a way to make your content on the web "licensed" in a way where it is only free for human consumption? I.e. effectively making the use of AI crawlers pirating, thus subject to the same kind of penalties here?

No. Neither legally or technically possible.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#98
post #62

Earlier quoted context omitted.

The blogger’s content was freely available, this fine is for piracy.

This is not a fine, it's a settlement to recompense authors. More broadly, I think that's a goofy argument. The books were "freely available" too. Just because something is out there, doesn't necessarily mean you can use it however you want, and that's the crux of the debate.

It's not the crux of this case. This is a settlement based on the judge's ruling that they books had been illegally downloaded. The same judge said that the training itself was not the problem – it was downloading the pirated books. It will be tough to argue that loading a public website is an illegal download.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#99
post #8

To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.

> It’s important in the fair use assessment to understand that the training itself is fair use, I think that this is a distinction many people miss. If you take all the works of Shakespeare, and reduce it to tokens and vectors is it Shakespeare or is it factual information about Shakespeare? It is the latter, and as much as organizations like the MLB might want to be able to copyright a fact you simply cannot do that…

> And what about all the other stuff that LLM's spit out? Who owns that. Well at present, no one. If you train a monkey or an elephant to paint, you cant copyright that work because they aren't human, and neither is an LLM.

This seems too cute by half, courts are generally far more common sense than that in applying the law.

This is like saying using `rails generate model:example` results in a bunch of code that isn't yours, because the tool generated it according to your specifications.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#100
post #8

To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.

> It’s important in the fair use assessment to understand that the training itself is fair use, I think that this is a distinction many people miss. If you take all the works of Shakespeare, and reduce it to tokens and vectors is it Shakespeare or is it factual information about Shakespeare? It is the latter, and as much as organizations like the MLB might want to be able to copyright a fact you simply cannot do that…

I mean, sort of. The issue is that the compression is novel. So anything post tokenization could arguably be considered value add and not necessarily derivative work.
Post reply on HN