Live data from Hacker News

Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

businessinsider.com

221–230 of 686 posts

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#221

I'm not seeing how this is fair use in either case. Someone correct me if I am wrong but aren't these works being digitized and transformed in a way to make a profit off of the information that is included in these works? It would be one thing for an individual to make person use of one or more books, but you got to have some special blindness not to see that a for-profit company's use of this information to improve…

There is another case where companies slurped up all of the internet and profited off the information, that makes a good comparison - search engines.

Judges consider a four factor when examining fair use[1]. For search engines,

1) The use is transformative, as a tool to find content is very different purpose than the content itself.

2) Nature of the original work runs the full gamut, so search engines don't get points for only consuming factual data, but it was all publicly viewable by anyone as opposed to books which require payment.

3) The search engine store significant portions of the work in the index, but it only redistributes small portions.

4) Search engines, as original devised, don't compete with the original, in fact they can improve potential market of the original by helping more people find them. This has changed over time though, and search engines are increasingly competing with the content they index, and intentionally trying to show the information that people want on the search page itself.

So traditional search which was transformative, only republished small amounts of the originals, and didn't compete with the originals fell firmly on the side of fair use.

Google News and Books on the other hand weren't so clear cut, as they were showing larger portions of the works and were competing with the originals. They had to make changes to those products as a result of lawsuits.

So now lets look at LLMs:

1) LLM are absolutely transformative. Generating new text at users request is a very different purpose and character from the original works.

2) Again runs the full gamut (setting aside the clear copyright infringement downloading of illegally distributed books which is a separate issue)

3) For training purposes, LLMs don't typically preserve entire works, so the model is in a better place legally than a search index, which has precedent that storing entire works privately can be fair use depending on the other factors. For inference, even though they are less likely to reproduce the originals in their outputs than search engines, there are failure cases where an LLM over-trained on a work, and a significant amount the original can be reproduced.

4) LLMs have tons of uses some of which complement the original works and some of which compete directly with them. Because of this, it is likely that whether LLMs are fair use will depend on how they are being used - eg ignore the LLM altogether and consider solely the output and whether it would be infringing if a human created it.

This case was solely about whether training on books is fair use, and did not consider any uses of the LLM. Because LLMs are a very transformative use, and because they don't store original verbatim, it weighs strongly as being fair use.

I think the real problems that LLMs face will be in factors 3 and 4, which is very much context specific. The judge himself said that the plaintiffs are free to file additional lawsuits if they believe the LLM outputs duplicate the original works.

[1] https://fairuse.stanford.edu/overview/fair-use/four-factors/

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#222

Earlier quoted context omitted.

Just downloading them is of course cheaper, but it is worth pointing out that, as the article states, they did also buy legitimate copies of millions of books. (This includes all the books involved in the lawsuit.) Based on the judgement itself, Anthropic appears to train only on the books legitimately acquired. Used books are quite cheap, after all, and can be bought in bulk.

Buying a book is not license to re-sell that content for your own profit. I can't buy a copy of your book, make a million Xeroxes of it and sell those. The license you get when you buy a book is for a single use, not a license to do what ever you want with the contents of that book.

What are you on about - the judge has literally said this was not resell, and is transformative and fair use.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#223

If you own a book, it should be legal for your computer to take a picture of it. I honestly feel bad for some of these AI companies because the rules around copyright are changing just to target them. I don't owe copyright to every book I read because I may subconsciously incorporate their ideas into my future work.

The difference here is that an LLM is a mechanical process. It may not be deterministic (at least, in a way that my brain understands determinism), but it's still a machine.

What you're proposing is considering LLMs to be equal to humans when considering how original works are created. You could make the argument that LLM training data is no different from a human "training" themself over a lifetime of consuming content, but that's a philosophical argument that is at odds with our current legal understanding of copyright law.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#224

Earlier quoted context omitted.

Saying that piracy isn't copyright violation is an RMS talking point. It's not worth trying to ask why because the answer will be RMS said so and will not be backed by the common usage of the word.

You legitimately have it completely backwards. The word "piracy" was coopted to put a more severe spin on copyright violation. As a result, it became "the common usage of the word". But that was by design. And it's worth pushing back on.

Sweden has a political party called "The Pirate Party"(1), and "The Pirate Bay" is Swedish so I think a couple of Swedes memeing before it was cool has a significant impact on making the name stick but also taking the seriousness out of it.

1: https://piratpartiet.se/en/

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#225
post #57

Earlier quoted context omitted.

With Claude, people are paying Anthropic to access answers that are generated from pirated books, without the authors permission, credit, or compensation.

There is no copyright on knowledge. If it outputs parts of the book verbatim then that's a different story.

>If it outputs parts of the book verbatim then that's a different story.

But it does...

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#226
post #2

Anthropic's cofounder, Ben Mann, downloaded million copies of books from Library Genesis in 2021, fully aware that the material was pirated. Stealing is stealing. Let's stop with the double standards.

[flagged]

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#227

Earlier quoted context omitted.

Saying that piracy isn't copyright violation is an RMS talking point. It's not worth trying to ask why because the answer will be RMS said so and will not be backed by the common usage of the word.

You legitimately have it completely backwards. The word "piracy" was coopted to put a more severe spin on copyright violation. As a result, it became "the common usage of the word". But that was by design. And it's worth pushing back on.

I don't have it backwards. Language evolved, and piracy got a new definition. It's even in the dictionary. Trying to redefine words like this is futile and avoiding certain words or replacing them with others is a quirk that RMS has.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#228

If you own a book, it should be legal for your computer to take a picture of it. I honestly feel bad for some of these AI companies because the rules around copyright are changing just to target them. I don't owe copyright to every book I read because I may subconsciously incorporate their ideas into my future work.

The difference here is that an LLM is a mechanical process. It may not be deterministic (at least, in a way that my brain understands determinism), but it's still a machine. What you're proposing is considering LLMs to be equal to humans when considering how original works are created. You could make the argument that LLM training data is no different from a human "training" themself over a lifetime of consuming cont…

[deleted]

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#230

  Alsup detailed Anthropic's training process with books: The OpenAI rival 
  spent "many millions of dollars" buying used print books, which the 
  company or its vendors then stripped of their bindings, cut the pages, 
  and scanned into digital files.
I've noticed an increase in used book prices in the recent past and now wonder if there is an LLM effect in the market.
Post reply on HN