Live data from Hacker News

A federal judge sides with Anthropic in lawsuit over training AI on books

techcrunch.com

211–220 of 222 posts

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#211
post #47

Earlier quoted context omitted.

Fair use overrides licensing

Fair use "overrides" licensing in the sense that one doesn't need a copyright license if fair use applies. But fair use itself isn't a shield against breach of contract. If you sign a license contract saying you won't train on the thing you've licensed, the licensor still has remedies for breach of contract, just not remedies for copyright infringement (assuming the act is fair use).

I am not going to sign a contract at the bookstore. Anyone who tries to get me to sign a contract at the bookstore is just going to lose book sales. IIRC the case involved Anthropic literally feeding physical books into scanners. Your proposed solution sounds like its just going to make books worse, not AI better.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#212
post #47

Earlier quoted context omitted.

Fair use "overrides" licensing in the sense that one doesn't need a copyright license if fair use applies. But fair use itself isn't a shield against breach of contract. If you sign a license contract saying you won't train on the thing you've licensed, the licensor still has remedies for breach of contract, just not remedies for copyright infringement (assuming the act is fair use).

I am not going to sign a contract at the bookstore. Anyone who tries to get me to sign a contract at the bookstore is just going to lose book sales. IIRC the case involved Anthropic literally feeding physical books into scanners. Your proposed solution sounds like its just going to make books worse, not AI better.

I suspect IP like text is going to follow the college virtual textbook model where DRMed software is needed to access it and physical copies won't exist. Maybe some HDCP-like protection to stop screen scraping.

To access them, institutions do have to sign contracts, along with abiding by licensing terms.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#213
post #53

Earlier quoted context omitted.

It might be different if you are a commercial product which couldn’t have been created without incorporating the contents of all those books. Humans, animals, hardware and software are treated differently by law because they have different constraints and capabilities.

But a commercial product is reaching parity with human capability. Let's be real, Humans have special treatment (more special than animals as we can eat and slaughter animals but not other humans) because WE created the law to serve humans. So in terms of being fair across the board LLMs are no different. But there's no harm in giving ourselves special treatment.

>So in terms of being fair across the board LLMs are no different

Why should "fair" factor into it? The LLMs are not humans, thus they have no rights, and treating them fairly shouldn't come into it. Stop anthropomorphizing linear algebra ffs.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#214

The reason I made books3 was to help force a decision on this issue. I’m happy to see that it’s settled, and that it’s legal for robots to read books. It’s also proof that an individual scientist can still change the world, in some small way. Believe in yourself and just focus on your work, even if the work is controversial. (I’m late to the thread, so ~nobody will see this. But it’s the culmination of about five yea…

I think you have indirectly done a disservice to the artistic community.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#215
post #182

Earlier quoted context omitted.

Its already established it can recite whole Hairy Potter and Carmacks Fast Inverse word for word. Just because it uses fancy compression doesnt mean its not a copy.

It can recite something like 80% of Harry Potter with carefully crafted prompts . If you take half a sentence from Harry Potter then tell the LLM to predict the rest it will complete it. That's what they did in that study you're referring to. It's not even remotely the same thing as "can recite whole Harry Potter." If you ask an LLM to regurgitate Harry Potter it won't be able to do so because that's not how they wor…

>predict

unpack, unless you are going to convince me LLMs are predicting '0x5f3759df' :). Lossy compression is still compression.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#216
post #195

Earlier quoted context omitted.

A judge already ruled that models themselves don't constitute copyright infringement in Kadrey v. Meta Platforms, Inc. ( https://casetext.com/case/kadrey-v-meta-platforms-inc ). The EFF has a good summary about it: > the court dismissed “nonsensical” claims that Meta’s LLaMA models are themselves infringing derivative works. See: https://www.eff.org/deeplinks/2025/02/copyright-and-ai-cases...

Time to overfit on some books and publicize them as a libgen mirror.

I think this could lead to interesting results outside the legalities.

Imagine you're getting it to spit out lord of the rings, but midway through you inject into the output 'Suddenly, the ring split in two. No longer one ring to rule them all, but two!'.

You then let the model write the rest of the story!

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#217
post #195

Earlier quoted context omitted.

Time to overfit on some books and publicize them as a libgen mirror.

I think this could lead to interesting results outside the legalities. Imagine you're getting it to spit out lord of the rings, but midway through you inject into the output 'Suddenly, the ring split in two. No longer one ring to rule them all, but two!'. You then let the model write the rest of the story!

I'm sure many people have imaged this - supposing that LLMs, while making no great strides towards AGI, consciousness, or any of that, nonetheless keep getting better and better at what they do now. Imagine a decade or two of steady improvements, throw in at least a couple of major breakthroughs. Much longer context by a few orders of magnitude. Much better quality, in terms of tone, consistency, hallucinations.

Maybe we'll actually be able to say things like: write me a trilogy in the style of Lord of the Rings but with these changes:

* Make it scifi

* Add more female characters with greater depth

* At least five rings

* Hobbits are the bad guys

... Or whatever, specifying a version of the story tailored to your intersts, and that you would get out really high quality results, similar in quality to the source materials.

Imagine you could do the same with movies, games, music.

I'm not trying to assign a value judgement here. There's good and bad sides. However, this reality is becoming easier to imagine with each new model released.

For sure, anyone who is a writer or artist will see this as bad. But perhaps our whole concept of what art is will become more fluid and personalized.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#218

One aspect of this ruling [1] that I find concerning: on pages 7 and 11-12, it concedes that the LLM does substantially "memorize" copyrighted works, but rules that this doesn't violate the author's copyright because Anthropic has server-side filtering to avoid reproducing memorized text. (Alsup compares this to Google Books, which has server-side searchable full-text copies of copyrighted books, but only allows user…

> One aspect of this ruling [1] that I find concerning: on pages 7 and 11-12, it concedes that the LLM does substantially "memorize" copyrighted works,

No, it doesn't. The order assumes that because it is an order on summary judgement, and the legal standard for such an order is that it must assume the least favorable position for the party for whom summart judgement is granted on every material contested issue of fact. Since it is a ruling for the defendant (Anthropic), it must be what the judge finds law demands when assuming all contested issues of fact are resolved in favor of the claims of the plaintiffs (the authors).

> but rules that this doesn't violate the author's copyright because Anthropic has server-side filtering to avoid reproducing memorized text.

No, it doesn't do that, either. It simply notes for clarity that the plaintiffs do not allege that that an infringement is created by the outputs for the reason you describe; the ruling does not in any way suggest that has any bearing on its findings as regards whether training the model infringes, it simply points out that that separate potential source of infringement is not at issue.

> Does this imply that distributing open-weights models such as Llama is copyright infringemen

No, it does not. At most, it implies, given the reason that rhe plaintiffs have not done so in this case, that the same plaintiffs might have alleged (without commenting at all as to whether they would prevail) that providing a hosted online service without filtering would constitute contributory infringement if that was what Anthropic did (which it isn’t) and if there was actual infringement committed by the users of the service.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#219
post #47

Earlier quoted context omitted.

Fair use "overrides" licensing in the sense that one doesn't need a copyright license if fair use applies. But fair use itself isn't a shield against breach of contract. If you sign a license contract saying you won't train on the thing you've licensed, the licensor still has remedies for breach of contract, just not remedies for copyright infringement (assuming the act is fair use).

I am not going to sign a contract at the bookstore. Anyone who tries to get me to sign a contract at the bookstore is just going to lose book sales. IIRC the case involved Anthropic literally feeding physical books into scanners. Your proposed solution sounds like its just going to make books worse, not AI better.

I'm not proposing any kind of solution, just stating what the law currently is. A book purchased at a store is a purchase; content obtained from online services like Bloomberg or LexisNexis is typically licensed; more and more of these license contracts include AI-focused restrictions.
Post reply on HN