Live data from Hacker News

Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

businessinsider.com

271–280 of 686 posts

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#271

Earlier quoted context omitted.

That's not a philosophical argument at odds with our current understanding of copyright law. That's exactly what this judge found copyright law currently is and it's quoted in the article being discussed.

Thanks for pointing that out. Obviously I hadn't read the whole article. That is an interesting determination the judge made: > Alsup ruled that Anthropic's use of copyrighted books to train its AI models was "exceedingly transformative" and qualified as fair use, a legal doctrine that allows certain uses of copyrighted works without the copyright owner's permission.

There are still questions: is an AI a 'user' in the copyright sense?

Or even, is an individual operating within the law as fair use, the same as a voracious all-consuming AI training bot consuming everything the same in spirit?

Consider a single person in a National Park, allowed to pick and eat berries, compared to bringing a combine harvester to take it all.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#272

If AI companies are allowed to use pirated material to create their products, does it mean that everyone can use pirated software to create products? Where is the line? Also please don't use word "learning", use "creating software using copyrighted materials". Also let's think together how can we prevent AI companies from using our work using technical measures if the law doesn't work?

It's abusive and wrong to try and prevent AI companies from using your works at all. The whole point of copyright is to ensure you're paid for your work. AI companies shouldn't pirate, but if they pay for your work, they should be able to use it however they please, including training an LLM on it. If that LLM reproduces your work, then the AI company is violating copyright, but if the LLM doesn't reproduce your work…

> The whole point of copyright is to ensure you're paid for your work.

No. The point of copyright is that the author gets to decide under what terms their works are copied. That's the essence of copyright. In many cases, authors will happily sell you a copy of their work, but they're under no obligation to do so. They can claim a copyright and then never release their work to the general public. That's perfectly within their rights, and they can sue to stop anybody from distributing copies.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#274

Earlier quoted context omitted.

Can you explain why? What makes them categorically different or at the very least why is "piracy" quantitatively worse than 'just' copyright violation?

Saying that piracy isn't copyright violation is an RMS talking point. It's not worth trying to ask why because the answer will be RMS said so and will not be backed by the common usage of the word.

> RMS

Referring to this? (Wikipedia's disambiguation page doesn't seem to have a more likely article.)

https://en.wikipedia.org/wiki/Richard_Stallman#Copyright_red...

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#275

Buying, scanning, and discarding was in my proposal to train under copyright restrictions. You are often allowed to nake a digital copy of a physical work you bought. There are tons of used, physical works thay would be good for training LLM's. They'd also be good for training OCR which could do many things, including improve book scanning for training. This could be reduced to a single act of book destruction per co…

[flagged]

That's true and was the distinction I was making. In my proposal, and maybe part of what Anthropic did, the digitized copies are used as training data for a new work, the model. That reduces the risk of legal rulings against using the copyrighted works.

From there, the cases would likely focus on whether that fits in established criteria for digitized copies, whether they're allowed in the training process itself, and the copyright status of the resulting model. Some countries allow all of that if you legally obtained the material in the first place. Also, they might factor whether it's for commercial use or not.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#276
post #250

From Vinge's "Rainbow's End": > In fact this business was the ultimate in deconstruction: First one and then the other would pull books off the racks and toss them into the shredder's maw. The maintenance labels made calm phrases of the horror: The raging maw was a "NaviCloud custom debinder." The fabric tunnel that stretched out behind it was a "camera tunnel...." The shredded fragments of books and magazine flew do…

[deleted]

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#277

Earlier quoted context omitted.

It is not wrong at all. The author decides what to do with their work. AI companies are rich and can simply buy the rights or hire people to create works. I could agree with exceptions for non-commercial activity like scientific research, but AI companies are made for extracting profits and not for doing research. > AI companies shouldn't pirate, but if they pay for your work, they should be able to use it however th…

If you reproduce the material from a work you've purchased then of course you're in violation of copyright, but that's not what an LLM does (and when it does I already conceded it's in violation and should be stopped). An LLM that doesn't "sell goods with movie characters" is not in violation. And the harm you describe is not a recognized harm. You don't own information, you own creative works in their entirety. If y…

To load a printed book into a computer one has to reproduce it in digital form without authorization. That's making a copy.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#278

The important parts: > Alsup ruled that Anthropic's use of copyrighted books to train its AI models was "exceedingly transformative" and qualified as fair use > "All Anthropic did was replace the print copies it had purchased for its central library with more convenient space-saving and searchable digital copies for its central library — without adding new copies, creating new works, or redistributing existing copies…

You skipped quotes about the other important side: > But Alsup drew a firm line when it came to piracy. > "Anthropic had no entitlement to use pirated copies for its central library," Alsup wrote. "Creating a permanent, general-purpose library was not itself a fair use excusing Anthropic's piracy." That is, he ruled that - buying, physically cutting up, physically digitizing books, and using them for training is fair…

From my understanding:

> pirating the books for their digital library is not fair use.

"Pirating" is a fuzzy word and has no real meaning. Specifically, I think this is the cruz:

> without adding new copies, creating new works, or redistributing existing copies

Essentially: downloading is fine, sharing/uploading up is not. Which makes sense. The assertion here is that Anthropic (from this line) did not distribute the files they downloaded.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#280
post #77

Earlier quoted context omitted.

Properly remixing the content so that it can be considered distinct would be fair use. You can't copyright a style, concept or idea. Also mostly this would be a civil lawsuit for "damages".

It might be legal in the US, but not in the rest of the world. The trial is scheduled for December 2025. That’s when a jury will decide how much Anthropic owes for copying and storing over seven million pirated books

You make some good points, this really is going to take some careful judgment and chances are it's too complex for an actual courtroom to yield an ideal outcome.

Now places like Flea markets have been known to have a counterfeit DVD or two.

And there is more than one way to compare to non-digital content.

Regular books and periodicals can be sold out and/or out-of-print, but digital versions do not have these same exact limitations.

A great deal of the time though, just the opposite occurs, and a surplus is printed that no one will ever read, and which will eventually be disposed of.

Newspapers are mainly in the extreme category where almost always a significant number of surplus copies are intentionally printed.

It's all part of the same publication, a huge portion of which no one has ever rightfully expected for every copy to earn anything at all, much less a return on every single copy making it back to the original creator.

Which is one reason why so much material is supported by ads. Even if you didn't pay a high enough price to cover the cost of printing, it was all paid for well before it got into your hands.

Digital copies which are going unread are something like that kind of surplus. If you save it from the bin you should be able to do whatever you want with it either way, scan it how you see fit.

You just can't say you wrote it. That's what copyright is supposed to be for.

Like at the flea market, when two different vendors are selling the same items but one has legitimately purchased them wholesale and the other vendor obtained theirs as the spoils of a stolen 18-wheeler.

How do you know which ones are the pirated items?

You can tell because the original owners of the pirated cargo suffered a definite loss, and have none of it remaining any more.

OTOH, with things like fake Nikes at the flea market, you can be confident they are counterfeit whether they were stolen from anybody in any way or not.

Post reply on HN