Live data from Hacker News

Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

businessinsider.com

241–250 of 686 posts

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#241

If you own a book, it should be legal for your computer to take a picture of it. I honestly feel bad for some of these AI companies because the rules around copyright are changing just to target them. I don't owe copyright to every book I read because I may subconsciously incorporate their ideas into my future work.

The core problem here is that copyright already doesn't actually follow any consistent logical reasoning. "Information wants to be free" and so on. So our own evaluation of whether anything is fair use or copyrighted or infringement thereof is always going to be exclusively dictated by whatever a judge's personal take on the pile of logical contradictions is. Remember, nominally, the sole purpose of copyright is not rooted in any notions of fairness or profitability or anything. It's specifically to incentivize innovation.

So what is the right interpretation of the law with regards to how AI is using it? What better incentivizes innovation? Do we let AI companies scan everything because AI is innovative? Or do we think letting AI vacuum up creative works to then stochastically regurgitate tiny (or not so tiny) slices of them at a time will hurt innovation elsewhere?

But obviously the real answer here is money. Copyright is powerful because monied interests want it to be. Now that copyright stands in the way of monied interests for perhaps the first time, we will see how dedicated we actually were to whatever justifications we've been seeing for DRM and copyright for the last several decades.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#242

Earlier quoted context omitted.

Just downloading them is of course cheaper, but it is worth pointing out that, as the article states, they did also buy legitimate copies of millions of books. (This includes all the books involved in the lawsuit.) Based on the judgement itself, Anthropic appears to train only on the books legitimately acquired. Used books are quite cheap, after all, and can be bought in bulk.

Buying a book is not license to re-sell that content for your own profit. I can't buy a copy of your book, make a million Xeroxes of it and sell those. The license you get when you buy a book is for a single use, not a license to do what ever you want with the contents of that book.

Yes, of course! In this case, the judge identified three separate instances of copying: (1) downloading books without authorisation to add to their internal library, (2) scanning legitimately purchased books to add to their internal library, and (3) taking data from their internal library for the purposes of training LLMs. The purchasing part is only relevant for (2) — there the judge ruled that this is fair use. This makes a lot of sense to me, since no additional copies were created (they destroyed the physical books after scanning), so this is just a single use, as you say. The judge also ruled that (3) is fair use, but for a different reason. (They declined to decide whether (1) is fair use at this point, deferring to a later trial.)

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#243

Earlier quoted context omitted.

They clearly were being digitized, but I think its a more philosophical discussion that we're only banging our heads against for the first time to say whether or not it is fair use. Simply, if the models can think then it is no different than a person reading many books and building something new from their learnings. Digitization is just memory. If the models cannot think then it is meaningless digital regurgitation…

> Simply, if the models can think then it is no different than a person reading many books and building something new from their learnings. No, that's fallacious. Using anthropomorphic words to describe a machine does not give it the same kinds of rights and affordances we give real people.

The judge did use some language that analogized the training with human learning. I don't read it as basing the legal judgement on anthropomorphizing the LLM though, but rather discussing whether it would be legal for a human to do the same thing, then it is legal for a human to use a computer to do so.

  First, Authors argue that using works to train Claude’s underlying LLMs was like using
  works to train any person to read and write, so Authors should be able to exclude Anthropic
  from this use (Opp. 16). But Authors cannot rightly exclude anyone from using their works for
  training or learning as such. Everyone reads texts, too, then writes new texts. They may need
  to pay for getting their hands on a text in the first instance. But to make anyone pay
  specifically for the use of a book each time they read it, each time they recall it from memory,
  each time they later draw upon it when writing new things in new ways would be unthinkable.
  For centuries, we have read and re-read books. We have admired, memorized, and internalized
  their sweeping themes, their substantive points, and their stylistic solutions to recurring writing
  problems.

  ...

  In short, the purpose and character of using copyrighted works to train LLMs to generate
  new text was quintessentially transformative. Like any reader aspiring to be a writer,
  Anthropic’s LLMs trained upon works not to race ahead and replicate or supplant them — but
  to turn a hard corner and create something different. If this training process reasonably
  required making copies within the LLM or otherwise, those copies were engaged in a
  transformative use.
[1] https://authorsguild.org/app/uploads/2025/06/gov.uscourts.ca...

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#244

Earlier quoted context omitted.

[flagged]

I think I have worded my question wrong. I asked about not about how AI affects the financials of these smaller artists, but their wellbeing in general. There are many small artists who do this not for money, but for fun and have their renowned styles. Even their styles are ripped off by these generative AI companies and turned into a slot machine to earn money for themselves. These artists didn't consent to that, an…

(1) You can't copyright an art style. That's not a thing.

(2) Once you make something publicly available, anyone can learn from it. No consent necessary.

(3) Being upset does not grant you special privileges under the law.

(4) If you don't like the idea of paying for AI art, free software is both plentiful and competitive with just about anything proprietary.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#245

If you own a book, it should be legal for your computer to take a picture of it. I honestly feel bad for some of these AI companies because the rules around copyright are changing just to target them. I don't owe copyright to every book I read because I may subconsciously incorporate their ideas into my future work.

Something missed in arguments such as these is that in measuring fair use there's a consideration of impact on the potential market for a rightsholder's present and future works. In other words, can it be proven that what you are doing is meaningfully depriving the author of future income.

Now, in theory, you learning from an author's works and competing with them in the same market could meaningfully deprive them of income, but it's a very difficult argument to prove.

On the other hand, with AI companies it's an easier argument to make. If Anthropic trained on all of your books (which is somewhat likely if you're a fairly popular author) and you saw a substantial loss of income after the release of one of their better models (presumably because people are just using the LLM to write their own stories rather than buy your stuff), then it's a little bit easier to connect the dots. A company used your works to build a machine that competes with you, which arguably violates the fair use principle.

Gets to the very principle of copyright, which is that you shouldn't have to compete against "yourself" because someone copied you.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#246
If AI companies are allowed to use pirated material to create their products, does it mean that everyone can use pirated software to create products? Where is the line?

Also please don't use word "learning", use "creating software using copyrighted materials".

Also let's think together how can we prevent AI companies from using our work using technical measures if the law doesn't work?

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#247

Earlier quoted context omitted.

You skipped quotes about the other important side: > But Alsup drew a firm line when it came to piracy. > "Anthropic had no entitlement to use pirated copies for its central library," Alsup wrote. "Creating a permanent, general-purpose library was not itself a fair use excusing Anthropic's piracy." That is, he ruled that - buying, physically cutting up, physically digitizing books, and using them for training is fair…

So all they have to do is go and buy a copy of each book they pirated. They will have ceased and desisted.

I'm trying to find the quote, but I'm pretty sure the judge specifically said that going and buying the book after the fact won't absolve them of liability. He said that for the books they pirated they broke the law and should stand trial for that and they cannot go back and un-break in by buying a copy now.

Found it: https://www.nbcnews.com/tech/tech-news/federal-judge-rules-c...

> “That Anthropic later bought a copy of a book it earlier stole off the internet will not absolve it of liability for the theft,” [Judge] Alsup wrote, “but it may affect the extent of statutory damages.”

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#249

Apparently it's a common business practice. Spotify (even though I can't find any proof) seems to have build their software and business on pirated music. There is some more in this Article [0]. https://torrentfreak.com/spotifys-beta-used-pirate-mp3-files... Funky quote: > Rumors that early versions of Spotify used ‘pirate’ MP3s have been floating around the Internet for years. People who had access to the service in…

There's plenty of startups gone legitimate. Society underestimates the chasm that exists between an idea and raising sufficient capital to act on those ideas. Plenty of people have ideas. We only really see those that successfully cross it. Small things EULA breaches, consumer licenses being used commercially for example.

Uber

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#250
From Vinge's "Rainbow's End":

> In fact this business was the ultimate in deconstruction: First one and then the other would pull books off the racks and toss them into the shredder's maw. The maintenance labels made calm phrases of the horror: The raging maw was a "NaviCloud custom debinder." The fabric tunnel that stretched out behind it was a "camera tunnel...." The shredded fragments of books and magazine flew down the tunnel like leaves in tornado, twisting and tumbling. The inside of the fabric was stitched with thousands of tiny cameras. The shreds were being photographed again and again, from every angle and orientation, till finally the torn leaves dropped into a bin just in front of Robert. Rescued data. BRRRRAP! The monster advanced another foot into the stacks, leaving another foot of empty shelves behind it.

Post reply on HN