Live data from Hacker News

Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

apnews.com

471–480 of 654 posts

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#471

Earlier quoted context omitted.

>The class counsel's unreimbursed litigation expenses were $2.6m. In what sane state does it even get that high?

I've seen some YouTubes where lawyers were complaining about high bill rates and showing actual bills. One large firm billed the senior lawyers at $2500/hour and even the paralegals were billed at $600/hr.

That could pay for a lot of tokens.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#472

Earlier quoted context omitted.

This is not backed up by any evidence. Ripping CDs was never illegal. The DMCA made the circumvention of an effective copyright protection mechanism illegal, which made ripping DVDs and Blu-rays a crime. But that's separate from copyright itself. The RIAA sued Napster users not because they were converting files, but because they were obtaining them from others without a license.

There is a lot of evidence. I lived through it. Every family with children and an internet connection or MP3 player was terrified of getting ruined suddenly via a letter. It was in the news every day about some other grandpa or single mother losing their house. Ripping CDs was long illegal. Perhaps the Librarian of Congress made an exception. Now they hid everything behind a CCB that is like Arbitration so we will ne…

You're confused about what was actually illegal and what the industry wanted you to believe was illegal. They didn't want to take anybody to court for actually ripping a CD because they didn't want to lose and have the precedent set like it was in Sony versus Betamax. This isn't "bigger fish to fry". This is "terrified of the precedent".

> The dispute arises from a suit the RIAA filed against a man in Arizona who bought CDs, copied them into his computer as MP3 files, and then put them into a shared folder that other people could access through Kazaa, a computer program for sharing music. He's being sued for that last part.

Then they state what they wish were true:

> But according to Marc Fisher, legal documents and some statements by industry officials make it clear that the industry regards the simple act of copying a CD onto your computer or your iPod as illegal.

But just because they wished it to be did not make it so. Trillion-dollar companies have provided end-users with software to rip CDs (including iTunes), and there's never been a court case over it.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#473

I do recall that a "famous redditor" was driven to suicide for making works available and he wasn't even making money for it.

He wasn’t just some famous redditor. Aaron Swartz helped create Reddit and invented RSS. If I put my conspiracy theory hat one and I always get piled on for this theory in other online communities but I think it could be possible. The theory is I think Aaron found some very dark stuff while exploring the MIT private networks, things that he was not supposed to see and could be very damaging to a lot people if they we…

Your theory is that he found "very dark stuff" that is accessible to any MIT student? Schwartz didn't "hack" anything, he connected to their student network and downloaded articles that they had access to but the public didn't.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#474

Earlier quoted context omitted.

> it did enable a lot of good work to happen. How do we know that when we don't have a copy of the world without this regime? How much more and greater works could have been produced without such a repressive system? A really successful work becomes part of the culture, and remixing, derivatives and other modes of integrating cultural artifacts are prohibited. Why should we allow corporations to own our culture?

You’re arguing that freely remixing original work will give rise to greatness that’s even better than original work?

I think the idea is if "AI" can solve math proofs that humans haven't for a century then if "AI" freestyles stolen art and literature then it might create something as good if not better because of resources and processing power

There might be something to that logic but art and literature doesn't obey rules like math and copyright exists to protect creators

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#475

Earlier quoted context omitted.

Why is it different other than, "just cause?" No one seems to have actual reasoning to back it up while it feels very similar the other way around, is human brains and neural nets (notwithstanding that they're both called neurons) seem to learn similarly and can act on similar classes of problems like language and mathematics.

are you asking what the diff is between a human and an LLM?

I asked why, not what.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#476

Earlier quoted context omitted.

Why is it different other than, "just cause?" No one seems to have actual reasoning to back it up while it feels very similar the other way around, is human brains and neural nets (notwithstanding that they're both called neurons) seem to learn similarly and can act on similar classes of problems like language and mathematics.

What's the difference between humans and LLMs? The difference is that the creators of laws are humans, and the purpose of laws is for humans. Authors write books with the expectation that they will be read by humans, and copyright law was written with unstated assumptions, like the fact that books have an effect on a person's mind after reading it. IMO the spirit of the law would prohibit LLMs from training, and the…

Laws are not only for humans, there are laws for bots as well, like anti spam, but that's besides the point because in reality behind LLMs there are humans and so humans still control them, therefore laws target them too, and now it looks like humans using LLMs to train via ingestion of books is deemed fair use.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#477

Earlier quoted context omitted.

Yes exactly. That's basically what an emulator for a games console is for example, a reimplementation of the original.

So a console game loses its copyright if you emulate it?

A game is a specific work to be copied so no, but the system it runs on can still be without copyright.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#478
post #453

Earlier quoted context omitted.

Whatever "training" is, if you can't persuade the machine to spit substantially the same text back out verbatim, it's clearly not something that falls under copy right law either, because there's no copy. Yes, for some texts that's possible. But for the vast majority, it is not.

Can you cite any information on this not being possible for the vast majority? Or is it simply that the correct prompt hasn't been written for all possible cases? I also fail to see the difference if logic/harnessing is added around a vector database that can output the complete corpus, but simply is instructed not to. It very clearly is still compressing the information into the vector weights, and then recovering t…

Information entropy. The amount of data an LLM ingests cannot be compressed to the size of the weights even at maximum theoretical compression.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#479

A one time payment like 1.5B doesn’t do anything. There needs to be a royalty payment based on if the AI regurgitates existing ideas. That is probably the correct way to legislate this. If anything a human does can instantly be copied by an LLM, and then sent to all its subscribers, things need to change

[deleted]

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#480

Earlier quoted context omitted.

There's also a difference between an MP3 and a FLAC. Again, ask DJs how well they're getting away on that distinction.

That’s not the legal criterion that’s used. Using a different codec is different that using the idea of a book to write your own book.

the "codec" is not really the point.

playing an MP3 at a venue, streaming it or distributing it is a copyrighted act because, despite not being a verbatim copy of the original material, it is capable of producing a nearly-verbatim version of that intellectual property well enough that most people won't be able to notice the difference.

similarly, as has been shown (by numerous publishers and authors), LLMs are capable of producing nearly-verbatim versions of the texts they have been trained on, to a well enough quality that most people won't be able to notice the difference.

the fact that an MP3 cannot "paraphrase" or "summarize" the audio data is not what makes it copyrighted, and neither does the ability of an LLM to "paraphrase" or "summarize" the textual data it's been trained on, make it any less intellectual property theft

the motivation for the audio case is the sense that the listener will not care whether the DJ plays an MP3 (they didn't pay for) or plays the original record (they would have paid for).

similarly for the lossily compressed text engine aka LLM's case, many people will not care whether they get this textual information paraphrased or nearly verbatim from an LLM trained on pirated books, or the original books.

the fact that an LLM also has the ability to paraphrase or summarize the pirated textual information it's been trained on, doesn't really matter if it's also capable of producing nearly verbatim copies of (parts of) those texts.

to underline this point even more, we know that MP3s (and more modern and much more efficient codecs like OPUS, after that) have been psycho-acoustically optimized to store exactly the least amount of data that will get "the point" of that music across to the listener, to the extent that they do not need the original recording any more. this is the stated goal of lossy compressed audio, after all. well, it also happens to be the (pretty much stated) goal of LLM companies, to store exactly the least amount of data that will get the point of that text to the reader. and it does tend to cause the readers to not really care about the original book any more.

having said all that, I don't mean to argue to lock it all up. I actually mean to argue that we should demand that Anthropic and Open AI release their weights data, and if anyone were to happen to break into them and steal that data, I would have exactly zero pity for that. because fair is fair.

Post reply on HN