Live data from Hacker News

Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

apnews.com

641–650 of 654 posts

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#641
post #613

Earlier quoted context omitted.

"I can use a lossy compression algorithm such that the original could never be recovered from the image I've produced, but that derived image would surely be under copyright." I've tried my best to show where I think you're wrong. I think all that's left is arguing over the exact definitions of "recoverable" and "irretrievable". As I said, the courts will have to decide that.

I have no clue what you're talking about tbh. I can use a lossy compression algorithm that is definitionally holding less information than the original while still be subject to the copyright of the original. This is just obviously true, converting a PNG to a JPEG does not invalidate the copyright on the PNG. You can also produce an imagine using a lossy compression algorithm that is not subject to the copyright of t…

> I have no clue what you're talking about tbh.

Point an LLM at the conversation and ask it to ELI5 the competing arguments.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#642

Earlier quoted context omitted.

In most cases it’s for the public good (even if a lot of the public has to be taken kicking and screaming), atleast the examples I listed above are. Uber and Tesla both revolutionized entire industries. So did Anthropic.

Is it? Uber has been a net loss for society. It increased overall VMT, especially in urban areas, and drove riders away from transit. On the whole this is both bad for cities and bad for city dwellers. We have more traffic and less transit, precisely the opposite of what we should be aiming for. If corporations can selectively ignore whatever laws they don't like, we won't have a functioning society anymore. Our soci…

Transit is a win only is large cities. Everywhere else, where most people in the US live (only 6% of the US population lives in dense, walkable urban environments), public transit is not a viable option. Uber is a huge win. Before Uber everyone that lived in a suburb or in a rural and semi rural area (~94% of the population)had to call an expensive limo that would cost hundreds of dollars to get anywhere if you didn’t have a car. That’s a thing of the past now.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#643
post #641

Earlier quoted context omitted.

I have no clue what you're talking about tbh. I can use a lossy compression algorithm that is definitionally holding less information than the original while still be subject to the copyright of the original. This is just obviously true, converting a PNG to a JPEG does not invalidate the copyright on the PNG. You can also produce an imagine using a lossy compression algorithm that is not subject to the copyright of t…

> I have no clue what you're talking about tbh. Point an LLM at the conversation and ask it to ELI5 the competing arguments.

Indeed, not sure how this person above can't comprehend any competing arguments and seem to say they "don't understand" over and over again.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#644

Earlier quoted context omitted.

Possibly. Sort of like how Disney remixed basically everything from existing fairy tales for decades then made it so nobody else could.

That’s not accurate. All the Disney interpretations of pre-existing fairy tales also have takes by others, whether books, theatrical adaptations, or even movies! Take Snow White as an example: https://en.wikipedia.org/wiki/Snow_White So I’d be curious to hear about a counter example.

I guess I haven't looked into it much, but the chilling effect of having Disney lawyers scrutinize your project would be enough to deter most people. The slew of content coming out after these properties enter public domain from Disney is telling. But it's also just the irony that Disney used other people's stories to build their empire but then fought extremely hard so that nobody could use theirs for as long as society would possibly hold out (which in my opinion was way beyond the point of reason).

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#645
post #584
post #519

Earlier quoted context omitted.

This is why the idea of being "Vogelfrei" or "lawless" was honestly a terrifying concept in the middle ages. They are neither bound by law, nor protected by law. A lawless man can be struck down with force without persecution by law, because they are lawless.

> A lawless man can be struck down with force without persecution by law, because they are lawless. Doesn't this define modern day police force theory?

No, because these days we are quite explicit that the law binds us all. Citizen or not, legally present or not, criminal or not, every human being within the borders of the US (or any modern country) is entitled to the full protection of the law.

When an American becomes an outlaw, it doesn't mean the same thing as it did back then. There is no legal way for the government to deny someone their rights.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#646

Earlier quoted context omitted.

So will you owe life long compensation for all the knowledge you got from books too? How about all the pirated books, music, movies, etc you consumed? When will you set up a life long payment plan to corporations that own these rights, because I have a bridge to sell you if you think any of this settlement will go to any of the people who created anything. I’m guessing you have some kind of imagined idea of some smal…

> So will you owe life long compensation for all the knowledge you got from books too? You're just falling into the trap of anthropomorphizing the phrase "training" in the context of LLMs, which is not the same things as what humans do. There is no evidence they are the same thing and there is nothing to support the notion that what an LLM does when it "trains" on a book is equivalent to a human reading it.

Ok, what about a human summarizing it or taking notes? What about a human indexing it for later searches? Does it matter if they use notecards or if they do it on a computer?

What if they retouch a photo you've taken as a political message? Does it matter if the do it with Sharpie, Photoshop, or by feeding it into an LLM?

When you publish, you give up some control over your work. Other people are allowed to do things with it and IMO it does not matter if it's in their head, on paper, or in a computer.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#647
post #528
post #518

Earlier quoted context omitted.

> spitting out verbatim text > had produced near-verbatim replicas

Don't get too hung up on the preciseness of the copy - the courts won't. I doubt that spitting out an existing article with a few adjectives changed would be considered transformative.

What's almost certainly going to happen is that the courts will decide the infringing party in that situation is the one who willfully caused the violation.

Since most users of LLMs are not using them to run around copyright and this copyright stuff is aggressively trained out of models or blocked whenever possible, the models themselves are still transformative works not intended to facilitate infringement.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#648
post #604

Earlier quoted context omitted.

I think it's occurred. That's why I wrote this in the very next paragraph: "Yes, for some texts that's possible." https://arxiv.org/abs/2601.02671 The point is that for most texts, it is not possible. It's not able to recall what I wrote on Geocities in 1995, even though there's a good chance it was trained on it.

No one is suing over what you wrote on Geocities in 1995...

No but it's still copyrighted content and subject to the same protections.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#649
post #641

Earlier quoted context omitted.

I have no clue what you're talking about tbh. I can use a lossy compression algorithm that is definitionally holding less information than the original while still be subject to the copyright of the original. This is just obviously true, converting a PNG to a JPEG does not invalidate the copyright on the PNG. You can also produce an imagine using a lossy compression algorithm that is not subject to the copyright of t…

> I have no clue what you're talking about tbh. Point an LLM at the conversation and ask it to ELI5 the competing arguments.

Why? I'm very obviously right. Entropy has nothing to do with copyright. Any LLM will answer the same.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#650

Earlier quoted context omitted.

There are several distinct differences, I think the most pertinent one is that you can create seperate instances of the LLM that ingested that media without reingesting the media (i.e. you can copy the LLM but you can't copy a human). Copright law is screwed up for multiple reasons but it makes sense to treat LLM as different than a human consuming the media, especially if the LLM is created for profit (e.g. movie ow…

I don't understand how the first pertinent one is relevant, when the LLM can't recreate verbatim at length an entire work anyway. It is literally unable to based on information entropy, it is not big enough to contain all of the bits of the training data at maximum theoretical compression.

But if a single person "can't recreate verbatim at length an entire work anyway" why does copyright law treat playing music to guests in your house differently to playing it to a classroom? I'm not saying that copyright law isn't stupid I'm saying that these distinctions make a difference and the fact that you can't copy humans is a pertinent distinction. LLMs would be treated very differently if like humans you had to spend the training costs of each model for each running instance of that model, the fact that models can be copied is a core feature of these models and makes them distinct from showing 'training data' to a human.
Post reply on HN