Live data from Hacker News

Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

apnews.com

501–510 of 654 posts

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#501

Earlier quoted context omitted.

That’s not the legal criterion that’s used. Using a different codec is different that using the idea of a book to write your own book.

the "codec" is not really the point. playing an MP3 at a venue, streaming it or distributing it is a copyrighted act because, despite not being a verbatim copy of the original material, it is capable of producing a nearly-verbatim version of that intellectual property well enough that most people won't be able to notice the difference. similarly, as has been shown (by numerous publishers and authors), LLMs are capabl…

> similarly, as has been shown (by numerous publishers and authors), LLMs are capable of producing nearly-verbatim versions of the texts they have been trained on, to a well enough quality that most people won't be able to notice the difference.

If that is true, you have a legal claim and can sue them. I doubt that’s true in the general case though.

The “does it hurt the original publisher” is a test for fair use BTW, just because you hurt the sales of someone doesn’t necessarily make it copyright infringement. That is only relevant if you try to defend using fair use (and it’s only part of the test that’s used to decide fair use).

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#502

Earlier quoted context omitted.

You’re arguing that freely remixing original work will give rise to greatness that’s even better than original work?

Yes? I don't understand how this is even a question, this is exactly how it worked throughout human history.

In what way is copyright preventing it today?

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#503
post #57

This is not even than a slap on the wrist. Publishers who negotiated this really fucked up writers. According to US federal law, pirating a single copyrighted work and gaining commercial advantage of it (which Anthropic 100% did) represents five years in prison and a $250,000 fine. But it gets worse: "Penalties for a copyright infringement conviction may increase if the defendant has previous similar convictions, mad…

That is a false statement. Gaining commercial advantage means selling pirated copies which Anthropic absolutely did not do, so none of your following statements are correct either.

It is an indisputable fact that Anthropic made money on models trained on those pirated books.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#504
I'm glad to see Anthropic's nose bloodied but I'm still very worried about the ruling in this case as I can already see the wheels turning as a way to limit fair use.

For context, the ruling is basically, "AI training is fair use but building a library of pirated books to train on is not". This is obviously because Judge Alsup does not want to put AI under a de-facto ban, but he wants AI companies to have to care about copyright... which in my opinion is self-contradictory, but let's go along with the (paraconsistent) logic.

If we insist that every prior act up to a fair use must be lawful, then this means that fair use is not a right, but a privilege that is purchased alongside the work itself. This opens the door to Oracle-level shenanigans: so long as every legal avenue to watch a work is encumbered by, say, a DeWitt clause[0], you cannot legally review the work. There are actually copyright cases hinging on this: Triller Fight Club sued H3H3 for reviewing a pirated stream of a Logan Paul fight that lasted 40 seconds and lost, for obvious reasons. This case smells like an accidental overturning of this.

Would I rather live in a world where robots[1] aren't allowed to read copyrighted books, or a world where copyright owners have veto rights over any and all critical commentary of their work? I would happily choose the former every time.

[0] A contractual clause that prohibits the recipient of a work from reviewing it without written permission of the owner.

[1] Mind uploads inclusive

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#505

Earlier quoted context omitted.

the "codec" is not really the point. playing an MP3 at a venue, streaming it or distributing it is a copyrighted act because, despite not being a verbatim copy of the original material, it is capable of producing a nearly-verbatim version of that intellectual property well enough that most people won't be able to notice the difference. similarly, as has been shown (by numerous publishers and authors), LLMs are capabl…

> similarly, as has been shown (by numerous publishers and authors), LLMs are capable of producing nearly-verbatim versions of the texts they have been trained on, to a well enough quality that most people won't be able to notice the difference. If that is true, you have a legal claim and can sue them. I doubt that’s true in the general case though. The “does it hurt the original publisher” is a test for fair use BTW…

It's a perfectly sensible interpretation of international copyright law.

I'm not sure if you're serious with the suggestion I could sue them. These are both US corporations, that justice system is pretty much in shambles in particular when it concerns corporations as big as these AI ones. You can dig your heels in the sand to defend that system, but you will also have to dig your head in the sand about why Sam Altman doesn't have a Disney "influenced" avatar, but one "inspired by" Studio Gibli.

And I'm not sure if you're familiar with the concept of "fair use" in the US as it "works" in practice, it's almost insulting, ask any music education youtuber.

Also even if it would work (which it very much doesn't), whether it "hurts the original publisher" is actually literally one of the criteria for considering something fair use or not. Look it up.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#506
A good example of the problem with this settlement:

>It appears that LLMs have already incorporated APOSD EDIT: The text of the book _A Philosophy of Software Design_ ENDEDIT (which would seem to be illegal, since it is copyrighted). For example, I have asked ChatGPT questions about APOSD and it seems to be able to answer.

https://groups.google.com/g/software-design-book/c/_wl1DciZZ...

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#507

A good example of the problem with this settlement: >It appears that LLMs have already incorporated APOSD EDIT: The text of the book _A Philosophy of Software Design_ ENDEDIT (which would seem to be illegal, since it is copyrighted). For example, I have asked ChatGPT questions about APOSD and it seems to be able to answer. https://groups.google.com/g/software-design-book/c/_wl1DciZZ...

I don't see:

- what's APOSD

- "it's illegal since it's copyrighted" makes no sense to me

- The settlement should be exactly to cover their licenses for training

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#508
post #404
post #271

Earlier quoted context omitted.

Ideas are not protected by copyright, nor are facts. You need to have a very specific and 'creative' / 'substantial' expression of an idea for copyright to apply. The output of an LLM can be easily be such, but usually not.

I'll link to a previous comment of mine: https://news.ycombinator.com/item?id=48968156 > You need to have a very specific and 'creative' / 'substantial' expression of an idea for copyright to apply. The output of an LLM can be easily be such, but usually not. This is incomplete with current US law. You need the above (the typical copyright qualifiers) AND evidence of substantial human involvement in the creation. Min…

Just to be clear, what you're referring to is the current US standard for whether a work is copywritable, not whether training on data and "regurgitating existing ideas" is fair-use. The latter is what the GP comment was about:

> There needs to be a royalty payment based on if the AI regurgitates existing ideas. That is probably the correct way to legislate this. If anything a human does can instantly be copied by an LLM, and then sent to all its subscribers, things need to change

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#509
post #293

A one time payment like 1.5B doesn’t do anything. There needs to be a royalty payment based on if the AI regurgitates existing ideas. That is probably the correct way to legislate this. If anything a human does can instantly be copied by an LLM, and then sent to all its subscribers, things need to change

This settlement has basically nothing to do with LLMs. At least not as far as the courts are concerned. Alsup ruled [0] that feeding a book into an LLM is transformative and counts as fair use. Especially when they purchased a physical copy of the book, scanned it, and destroyed the original. But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as lon…

Hey I just came up with this idea, I'm going to feed copyrighted books into my LLM that remembers them verbatim, and then people pay me to ask the LLM for complete copies of a book.

Wait, no, not verbatim. It transforms upper case into lower case and vice versa.

Post reply on HN