Live data from Hacker News

Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

apnews.com

601–610 of 654 posts

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#601

Earlier quoted context omitted.

Copyright is what stops someone from copy+pasting a book that took years to write, then selling it $1 cheaper than the original author on Amazon or whatever and making a margin 1 million percent higher than the original author. Imagine a society without copyright… only physically intensive jobs could make money because everything else would be pirated, ripped-off or free. Thus, only those who are financially independ…

That's oversimplifying things. Copyright doesn't actually stop me from pirating a book or an mp3 right now. Heck, I'll just download a book right now. Bam. Done. Some things are so difficult to keep from being pirated, such a photographs, that saying the copyright system protects photographers strikes me as a bit silly. It does protect some commercial photographers if a magazine wants to sell their photo sometimes, b…

> Also there are other systems that might protect an author's financials. Off the top of my head I imagine you could do a netflix model where every citizen pays some taxes to consume intellectual property like a utility. Then the goverment finds a way to measure what is being consumed and gives each author a share based on the rate of consumption.

We already have these - CD taxes, government grants funded by general taxes, GEMA in Germany, even TV licenses.

They all universally suck and are extremely unfair in who gets paid by them.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#602

Earlier quoted context omitted.

> it did enable a lot of good work to happen. How do we know that when we don't have a copy of the world without this regime? How much more and greater works could have been produced without such a repressive system? A really successful work becomes part of the culture, and remixing, derivatives and other modes of integrating cultural artifacts are prohibited. Why should we allow corporations to own our culture?

Yeah, I can't AB test against a world without copyright at all, but I think there's sufficient evidence to believe that a lot of stuff would never have gotten done without copyright to ensure it could be done gainfully. The importance of striking a balance between incentivising creation and enriching culture was why the original copyright term was dramatically shorter. The modern term of owners life + 80 years or wha…

> Yeah, I can't AB test against a world without copyright at all, but I think there's sufficient evidence to believe that a lot of stuff would never have gotten done without copyright to ensure it could be done gainfully.

The problem is that there is also a lot of stuff that never got done because of copyright. And the extend to which works that were funded by exploiting copyright would not have been funded in any other way is also questionable.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#603
post #452

Earlier quoted context omitted.

William Roper: "So, now you give the Devil the benefit of law!" Sir Thomas More: "Yes! What would you do? Cut a great road through the law to get after the Devil?" William Roper: "Yes, I’d cut down every law in England to do that!" Sir Thomas More: "Oh? And when the last law was down, and the Devil turned ’round on you, where would you hide, Roper, the laws all being flat? This country is planted thick with laws, fro…

That's a cool quote but utterly useless. You have a strong opinion and no argument.

[deleted]

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#604
post #453

Earlier quoted context omitted.

Whatever "training" is, if you can't persuade the machine to spit substantially the same text back out verbatim, it's clearly not something that falls under copy right law either, because there's no copy. Yes, for some texts that's possible. But for the vast majority, it is not.

> if you can't persuade the machine to spit substantially the same text back out verbatim That's exactly what they've done in a number of the lawsuits, so I'm not sure why you think that hasn't occurred.

I think it's occurred. That's why I wrote this in the very next paragraph: "Yes, for some texts that's possible."

https://arxiv.org/abs/2601.02671

The point is that for most texts, it is not possible. It's not able to recall what I wrote on Geocities in 1995, even though there's a good chance it was trained on it.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#605
post #524

Earlier quoted context omitted.

> but that derived image would surely be under copyright. I wouldn't bet on that. https://en.wikipedia.org/wiki/Campbell%27s_Soup_Cans

I don't know that this really challenges anything relating to compression.

Ok, reductio ad absurdum.

Here's a highly compressed representation of The Lord of The Rings (all three volumes):

1

Obviously, fidelity when uncompressing it is not great, but I can assure you it was lossily compressed from the original text. Is it infringing the original's copyright? I have to assume you'd agree that the answer is "no".

If I had compressed it by removing the letters x y and z, I'd agree with you that my "compressed" version is infringing.

So what we've got here is a spectrum with two ridiculous extremes, and a question: When has the artifact been compressed so heavily that it no longer infringes the copyright of the original?

I suggest "irretrievability" is a pretty good threshold for that question. Otherwise you're into "we know it infringes our copyright. Don't ask us to prove it, we just know it, ok?"

Given the sheer volume of text that an LLM gets trained on, and how small the output is, it seems obvious that 99% of it can no longer be recovered - the process is "lossy" to the point of irretrievability, and only a statistical smear is left behind. That's why I think only the copyright claims that can show infringement in court (Harpy Potter, et al.) have merit. And a court will still have to decide "how much is too much" but at least there's case law for that.

(Incidentally, I compressed the Mona Lisa to a single pixel. It was #3D3526).

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#606
post #503

Earlier quoted context omitted.

That is a false statement. Gaining commercial advantage means selling pirated copies which Anthropic absolutely did not do, so none of your following statements are correct either.

It is an indisputable fact that Anthropic made money on models trained on those pirated books.

It is an indisputable fact that "commercial advantage" in terms of copyright means selling copies. It is also an indisputable fact that Anthropic did npt do that.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#607
post #605

Earlier quoted context omitted.

I don't know that this really challenges anything relating to compression.

Ok, reductio ad absurdum. Here's a highly compressed representation of The Lord of The Rings (all three volumes): 1 Obviously, fidelity when uncompressing it is not great, but I can assure you it was lossily compressed from the original text. Is it infringing the original's copyright? I have to assume you'd agree that the answer is "no". If I had compressed it by removing the letters x y and z, I'd agree with you tha…

My entire point is that compression is irrelevant. A lossily compressed image can absolutely still be in breach of the original's copyright. Showing that some kind of compressed artifact may not be doesn't change that.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#608

Earlier quoted context omitted.

I am not sure what you're trying to say. Compressing an image is creating a derived work, it is still subject to the copyright of the original.

What AI is doing is not analogous to compression so even if you think they are they same, it legally and technically speaking is not subject to the copyright of the original.

Are you missing the comment that I'd responded to?

> Information entropy. The amount of data an LLM ingests cannot be compressed to the size of the weights even at maximum theoretical compression.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#609

Earlier quoted context omitted.

As far as I'm concerned, the courts are wrong, and training on ill gotten copyrighted material is not fair use. Given the clear value of highly trained LLMs, the investment they have taken on, and the amount of disruption to the existing economy they stand to make, in a just world, the people who created the training data deserve some level of compensation. I think, in the US, they are very afraid of falling behind C…

So will you owe life long compensation for all the knowledge you got from books too? How about all the pirated books, music, movies, etc you consumed? When will you set up a life long payment plan to corporations that own these rights, because I have a bridge to sell you if you think any of this settlement will go to any of the people who created anything. I’m guessing you have some kind of imagined idea of some smal…

I paid for my education, thank you very much. I'm still paying for it.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#610

Earlier quoted context omitted.

What very very similar ways would that be?

I guess they presume it requires on the good will of everyone to live peacefully without violating (non-existent) property laws over, say, robbing you in your sleep.

It doesn't require that at all, on the contrary. Why don't you educate yourself on these basic assumptions?
Post reply on HN