Live data from Hacker News

Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

apnews.com

331–340 of 654 posts

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#331
post #93

Did someone forget to consult with the MPAA and the RIAA on this one? This is a joke of an outcome. $3k per book. How much was it per song for Napster?

The RIAA typically asked for around $2-4 per song to settle without a lawsuit, which would come to a total of a few thousand because they generally only went after people sharing over a thousand songs. In the couple of few where the party would not agree to a settlement and the RIAA sued, they would pick about 15 of the thousand+ songs to sue over. Statutory damages are a minimum of $750 per infringed work, so the to…

How did the 200 million dollar lawsuits for one song come about then?

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#332

$3000/book for effectively pirating a book is shamefully low.

What do you mean pirating? They don't even distribute the originals, LLMs are not for replication, we already have copying and internet for that. Why would we use a multi-billion parameter model to copy text? If we wanted the originals it would be easier to find them free, pirate or pay, if we use LLMs it is because we want something ELSE. And caring about content rights in a world with limitless content and scarce a…

They mean pirating. Why is it suddenly hard to understand when an LLM company is involved?

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#333

Earlier quoted context omitted.

Real question, if an LLM shouldn't be able to remix someone's written work, why should a robot be able to build a chair that kinda looks like a chair a carpenter built that one time? The carpenter was a human, and humans are not a service. Why this distinction only for intellectual work?

For the same reason that you can make a similar-looking chair, but you can’t distribute a fuzzy copy of Star Wars. The char isn’t a copyrighted work.

"The char isn’t a copyrighted work."

An Eames chair is, we just have a really high bar for what is copyrightable in the physical world, and it seems pointlessly discriminatory.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#334

Earlier quoted context omitted.

There’s a difference between an abstract idea and the concrete thing. Regurgitating an idea is different than repeating the text verbatim. Ideas are protected by patents, not copyright.

So now that we have a magical paraphrasing machine, we can just run any copyrighted work through it to remove the copyright? Cool, I get a GPL version of Microsoft Office.

"I get a GPL version of Microsoft Office."

Is this not..Libre?

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#335

Earlier quoted context omitted.

It's pretty easy to validate that everything they're saying is accurate. https://www.history.com/this-day-in-history/september-8/riaa... > in practice the RIAA offered defendants the option of establishing a “Clean Slate” by destroying all of their illegally acquired files and paying a settlement of approximately $3 per illegal song. The two notable cases were: 1) https://en.wikipedia.org/wiki/Capitol_Records,_Inc._v…

Weird, what about this? https://w2.eff.org/IP/P2P/riaa_at_four.pdf

Could you also make the argument here instead of just linking a 25 page PDF?

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#336

Earlier quoted context omitted.

So now that we have a magical paraphrasing machine, we can just run any copyrighted work through it to remove the copyright? Cool, I get a GPL version of Microsoft Office.

"I get a GPL version of Microsoft Office." Is this not..Libre?

LibreOffice is a different product from Microsoft Office.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#337
post #308
post #293

Earlier quoted context omitted.

This settlement has basically nothing to do with LLMs. At least not as far as the courts are concerned. Alsup ruled [0] that feeding a book into an LLM is transformative and counts as fair use. Especially when they purchased a physical copy of the book, scanned it, and destroyed the original. But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as lon…

It's easier to ask for forgiveness than permission, right? It seems to be the modus operandi of corporations in general: they commit any kind of infringement they want and then later they go for a settlement with a value that's, of course, not too big for a company too big to fail. In the meantime, the average person or company gets shafted. In my opinion, we are one step away from AI companies capturing the entirety…

You do indeed appear to have a valid point. Many "chosen" companies, like Uber for example, appear to have broken numerous laws. Legal action against many such companies comes suspiciously slowly, where they have already obtained massive profits and value, before the possibility of being shut down comes. Then, when they are finally pulled into court, they have all kinds of money for the best lawyers and have already paid the right politicians (and others).

When the legal judgements for wrongdoing are finally handed out, they often come across as just an inconvenience or kind of tax, which is easily handled in comparison to the profits they've already made. Yet, if average Joe or persons not considered as being of "the right type" were to do such actions, they quickly get the full book thrown at them. Often, the full measure of legal punishment, where their company and life is or about nearly over.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#338
post #293

A one time payment like 1.5B doesn’t do anything. There needs to be a royalty payment based on if the AI regurgitates existing ideas. That is probably the correct way to legislate this. If anything a human does can instantly be copied by an LLM, and then sent to all its subscribers, things need to change

This settlement has basically nothing to do with LLMs. At least not as far as the courts are concerned. Alsup ruled [0] that feeding a book into an LLM is transformative and counts as fair use. Especially when they purchased a physical copy of the book, scanned it, and destroyed the original. But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as lon…

As far as I'm concerned, the courts are wrong, and training on ill gotten copyrighted material is not fair use. Given the clear value of highly trained LLMs, the investment they have taken on, and the amount of disruption to the existing economy they stand to make, in a just world, the people who created the training data deserve some level of compensation. I think, in the US, they are very afraid of falling behind China, who doesn't give a shit about intellectual property, but that doesn't mean we aren't crossing an ethical boundary, acting like them.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#339

Earlier quoted context omitted.

once you read the book, and learn from it and use the knowledge in your livelihood - you have to perpetually pay the author? Doesn't sound right.

Why is everyone talking about human analogies when LLMs are not humans?

Because the process is the same. When you read a book you don‘t save it as a brain file, you form memories from it. Some people can recall verbatim bits here and there but I have never met someone regurgitating a book word for word. And I‘m pretty sure I can not ask ChatGPT to output the first chapter of Moby Dick word for word. I think that would be, rightfully, considered copyright infringement.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#340

Earlier quoted context omitted.

There’s a difference between an abstract idea and the concrete thing. Regurgitating an idea is different than repeating the text verbatim. Ideas are protected by patents, not copyright.

So now that we have a magical paraphrasing machine, we can just run any copyrighted work through it to remove the copyright? Cool, I get a GPL version of Microsoft Office.

That’s what a brain is
Post reply on HN