Live data from Hacker News

Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

businessinsider.com

521–530 of 686 posts

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#521

Earlier quoted context omitted.

If you do that and reproduce the covers or the protected elements thereof, you should absolutely expect to be sued.

So for example, if the bookstore has a nice 4k surveillance camera and you have access to it because you work there, sitting at home and using it to look at the cover art on all the books on display is something you'd expect to be sued over?

Re-read my comment: "If you do that and reproduce the covers or the protected elements thereof"

This conversation becomes incredibly unenjoyable when you pull rhetorical techniques like completely ignoring the entirety of what I wrote.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#522

Earlier quoted context omitted.

>It's inherent in the nature of the test. The most important fair use factor is the effect on the market for the work, so if the use would be uneconomical without fair use then the effect on the market is negligible because the alternative would be that the use doesn't happen rather than that the author gets paid for it. No, that's not the most important factor. The transformative factor is the most important. Effect…

> No, that's not the most important factor. The transformative factor is the most important. It's a four factor test because all of the factors are relevant, but if the use has negligible effect on the market for the work then it's pretty hard to get anywhere with the others. For example, for cases like classroom use, even making verbatim copies of the entire work is often still fair use. Buying a separate copy for e…

>It's a four factor test because all of the factors are relevant, but if the use has negligible effect on the market for the work then it's pretty hard to get anywhere with the others. For example, for cases like classroom use, even making verbatim copies of the entire work is often still fair use. Buying a separate copy for each student to use for only a few minutes would make that use uneconomical.

All four factors are not equally relevant which is something described in pretty much every single fair use opinion. Educational uses are educational uses and considered fair because of their educational purpose (purpose is one of the factors), again, not because it's expensive. Maybe next time try googling or using ChatGPT "fair use educational".

>We're talking about the temporary copies they make during training. Those aren't being distributed to anyone else.

It's your argument. Not mine. You do not understand the market harm factor and it has nothing to do with Anthropic's transaction costs. That's just fully outright absolutely incorrect application of law.

>Making a copy of everything on the internet is a prerequisite to making a search engine. It's something you have to do as a step to making the index, which is the transformative step. Are you suggesting that doing the first step is illegal or what do you propose justifies it?

The transformative step is why it's a fair use, not the "market harm" (which you misunderstand) or the made up argument that it's "too expensive". In fact, I said this like every single turn in our conversation so it's a bit perplexing to me that you can now ask me "do you mean that it being transformative is what makes it legal" when that was my exact argument three times.

>Anything with unreasonably high transaction costs. Why is that ridiculous? It doesn't exempt any of the normal stuff like an individual person buying an individual book.

It's ridiculous because of the example I gave. Things being expensive is not a defense to copyright infringement and copyright law has no obligation to make expensive business models work. Copyright has an obligation to make transformative business models work because of the overall good they provide to society. Describing it as a "transaction cost" just kicks the can down the road even further and doesn't deal with the substance, either. They could have gone to the major publishers and licensed books from them. They didn't. That's generally who they are being sued by. When they are being sued by copyright owners in the fringe examples you pointed to, they will become relevant then.

>They need to get as many books as possible, with the platonic ideal being every book. Whether or not the ideal is feasible in practice, the question is whether it's socially beneficial to impose a situation with excessively high transaction costs in order to require something with only trivial benefit to authors (potentially selling one extra copy).

Lol dude, it was your example, not mine. They do not need every single book. They aren't being sued over every single book anyway, so it's totally besides the point.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#523

Earlier quoted context omitted.

> summarizing or reciting portions of the contents This absolutely falls under copyright law as I understand it (not a lawyer). E.g. the disclaimer that rolls before every NFL broadcast. The notice states that the broadcast is copyrighted and any unauthorized use, including pictures, descriptions, or accounts of the game, is prohibited. There is wiggle room for fair use by news organizations, critics, artists, etc.

I can say "you cannot read this comment for any purpose" but that doesn't supersede the law.

Btu it is ilelgal to rveerse enigneer tihs porprietary encrpytion algroithm I cerated and uesd to encrpyt tihs mesasge.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#524
post #50

Earlier quoted context omitted.

This is absurd. Remove all of the content from the training data that was pirated and what is the quality of the end product now?

That's the law. Please keep in mind, copyright is intended as a compromise between benefit to society and to the individual. A thought experiment, students pirating textbooks and applying that knowledge later on in their work?

Its the law (for now, very early on this in the process of deciding the law, untested, appealable, likely to be appealed and tested many times in many ways).

Meanwhile other cases have been less friendly to it being fair use, AI companies are already paying vast sums to publishers who presumably they wouldn’t if they felt confident it was “the law”, and on and on.

I don’t like arguing from “it’s the law”. A lot of law is terrible. What’s right? It’s clear to me that if AI gets good enough, as it nearly is now, it sucks a lot of profit away from creators. That is unbalanced. The AI doesn’t exist without the creators, the creators need to exist for our society to be great (we want new creative works, more if anything). Law tends to start conservatively based on historical precedent, and when a new technology comes along it often errs on letting it do some damage to avoid setting a bad precedent. In time it catches up as society gets a better view of things.

The right thing is likely not to let our creative class be decimated so a few tech companies become fantastically wealthy - in the long run, it’s the right thing even for the techies.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#525
post #310

Earlier quoted context omitted.

We're talking about a summary judgement issued that has not yet been appealed. That doesn't make it "settled." If by "what is stored and the manner which it is stored" is intended to signal model weights, I'm not sure what the argument is? The four factors of copyright in no way mention a storage medium for data, lossless or loss-y. (1) the purpose and character of the use, including whether such use is of a commerci…

The use is to train an AI model. A trillion parameter SOTA model is not substantially comprised of the one copyrighted piece. (If it was a Harry Potter model trained only on Harry Potter books this would be a different story). Embeddings are not copy paste. The last point about market impact would be where they make their argument but it's tenuous. It's not the primary use of AI models and built in prompts try to avo…

I bet it’s pretty easy to reproduce enough of Harry Potter from these models that any judge would see it as not fair use - you’d just have to prompt it in the right way. I’d bet a large sum that when this eventually shakes through the Supreme Court, it won’t be deemed fair use entirely, for the better of the world.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#526

Earlier quoted context omitted.

> the numerous media industry lawsuits against individuals that only mention downloading, I never saw any of these. All the cases I saw were related to people using torrents or other P2P software (which aren't just downloading). These might exist, but I haven't seen them. > It's a bit surprising that you can suddenly download copyrighted materials for personal use and it's kosher as long as you don't share them with…

I just checked first individual suit I could find, which was BMG v. Gonzalez. She used P2P, but the case was specifically about her downloading , not redistributing.

Most P2P tools work in a way where you cannot download without simultaneously uploading.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#527

The important parts: > Alsup ruled that Anthropic's use of copyrighted books to train its AI models was "exceedingly transformative" and qualified as fair use > "All Anthropic did was replace the print copies it had purchased for its central library with more convenient space-saving and searchable digital copies for its central library — without adding new copies, creating new works, or redistributing existing copies…

How times change .They wanted to lock up Aaron Schwartz for life for essentially doing the same thing Anthropic is doing.

Aaron Swartz wanted to provide the public with open access to paywalled journal articles, while Anthropic want to use other people's copyrighted material to train their own private models that they restrict access to via a paywall. It's wild (but unsurprising) that Aaron Swartz was prosecuted under the CFAA for this while Anthropic is allowed to become commercially successful

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#528
post #414

Earlier quoted context omitted.

These are two separate actions that Anthropic did: * They downloaded a massive online library of pirated books that someone else was distributing illegally. This was not fair use. * They then digitised a bunch of books that they physically owned copies of. This was fair use. This part of the ruling is pretty much existing law. If you have a physical book (or own a digital copy of a book), you can largely do what you…

The judge said they can train however I believe the judge did not make any ruling regarding model outputs

Thanks for the clarification!

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#529
post #481

Earlier quoted context omitted.

> Yes! But Suno is definitely not training models in their basement for fun. They are a private company selling music, using music made by humans to train their models, to replace human musicians and artists. We'll see what the courts say but that doesn't sound like fair use.

My understanding is that Suno does not sell music , but instead makes a tool for musicians to generate music and sells access to this tool . The law doesn't distinguish between basement and cloud – it's a service. You can sell access to the service without selling songs to consumers.

That's like arguing that a restaurant doesn't sell food because it sells the service of cooking it.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#530
post #481

Earlier quoted context omitted.

> Yes! But Suno is definitely not training models in their basement for fun. They are a private company selling music, using music made by humans to train their models, to replace human musicians and artists. We'll see what the courts say but that doesn't sound like fair use.

My understanding is that Suno does not sell music , but instead makes a tool for musicians to generate music and sells access to this tool . The law doesn't distinguish between basement and cloud – it's a service. You can sell access to the service without selling songs to consumers.

What does "fair use" even mean in a world where models can memorise and remix every book and song ever written? Are we erasing ownership?

The problem is, copyright law wasn't written for machines. It was written for humans who create things.

In the case of songs (or books, paintings, etc), only humans and companies can legally own copyright, a machine can't. If an AI-powered tool generates a song, there’s no author in the legal sense, unless the person using the tool claims authorship by saying they operated the tool.

So we're stuck in a grey zone: the input is human, the output is AI generated, and the law doesn't know what to do with that.

For me the real debate is: Do we need new rules for non-human creation?

Post reply on HN