Live data from Hacker News

Anthropic agrees to pay $1.5B to settle lawsuit with book authors

nytimes.com

101–110 of 761 posts

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#101
post #24
post #8

To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.

I guess they must delete all models since they acquired the source illegally and benefitted from it, right? Otherwise it just encourages others to keep going and pay the fines later.

[deleted]

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#103
post #55

Earlier quoted context omitted.

I wonder what Aaron Swartz would think if he lived to see the era of libgen.

He died (2013) after libgen was created (2008).

I had no idea libgen was that old, thanks!

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#104
> "The technology at issue was among the most transformative many of us will see in our lifetimes"

A judge making on a ruling based on his opinion of how transformative a technology will be doesn't inspire confidence. There's an equivocation on the word "transformative" here -- not just transformative in the fair use sense, but transformative as in world-changing, impactful, revolutionary. The latter shouldn't matter in a case like this.

> Companies and individuals who willfully infringe on copyright can face significantly higher damages — up to $150,000 per work

Settling for 2% is a steal.

> In June, the District Court issued a landmark ruling on A.I. development and copyright law, finding that Anthropic’s approach to training A.I. models constitutes fair use,” Aparna Sridhar, Anthropic’s deputy general counsel, said in a statement.

This is the highest-order bit, not the $1.5B in settlement. Anthropic's guilty of pirating.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#105
post #28

Earlier quoted context omitted.

This implies training models is some sort of right.

No, it implies that having the power to train AI models exclusively consolidated into a handful of extremely powerful companies is bad.

That's true. Those handful of companies shouldn't get to do it either.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#106
post #62

Earlier quoted context omitted.

The blogger’s content was freely available, this fine is for piracy.

This is not a fine, it's a settlement to recompense authors. More broadly, I think that's a goofy argument. The books were "freely available" too. Just because something is out there, doesn't necessarily mean you can use it however you want, and that's the crux of the debate.

But you can use copyrighted works for transformative works under the fair-use doctrine, and training was ruled to be fair use in the previous ruling.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#108

Earlier quoted context omitted.

Yeah but did he die before anybody actually knew about it?

Is lib still around anymore. I can't find any functioning urls

I believe that there's a reddit sub that keeps people up to date with what URLs are, or are not, functioning at any given point in time

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#110
post #8

To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.

> It’s important in the fair use assessment to understand that the training itself is fair use, I think that this is a distinction many people miss. If you take all the works of Shakespeare, and reduce it to tokens and vectors is it Shakespeare or is it factual information about Shakespeare? It is the latter, and as much as organizations like the MLB might want to be able to copyright a fact you simply cannot do that…

Yes. Someone on this post mentioned that switzerland allows downloading copyrightable material but not distributing them.

So things get even more dark because what becomes distribution can have a really vague definition and maybe the AI companies will only follow the law just barely, just for the sake of not getting hit with a lawsuit like this again. But I wonder if all this case did was maybe compensate the authors this one time. I doubt if we can see a meaningful change towards AI companies attitude's towards fair use/ essentially exploiting authors.

I feel like that they would try to use as much legalspeak as possible to extract as much from authors (legally) without compensating them which I feel is unethical but sadly the law doesn't work on ethics.

Post reply on HN