Live data from Hacker News

A federal judge sides with Anthropic in lawsuit over training AI on books

techcrunch.com

191–200 of 222 posts

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#191
post #182

Earlier quoted context omitted.

LLMs do not "make a copy of the entire book in its memory" so that specific question is kind of moot.

Its already established it can recite whole Hairy Potter and Carmacks Fast Inverse word for word. Just because it uses fancy compression doesnt mean its not a copy.

[deleted]

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#192

The reason I made books3 was to help force a decision on this issue. I’m happy to see that it’s settled, and that it’s legal for robots to read books. It’s also proof that an individual scientist can still change the world, in some small way. Believe in yourself and just focus on your work, even if the work is controversial. (I’m late to the thread, so ~nobody will see this. But it’s the culmination of about five yea…

FWIW, I see your comment. Also late to the thread though. This ruling is being watched at my office. I want to be a bit anonymous, but we've been doing a much more analogue version of some of these things for 75 years. With academics being our primary market. We've only had two legal issues in that time. Both settled out of court. But we walk a fine line.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#193
post #192

The reason I made books3 was to help force a decision on this issue. I’m happy to see that it’s settled, and that it’s legal for robots to read books. It’s also proof that an individual scientist can still change the world, in some small way. Believe in yourself and just focus on your work, even if the work is controversial. (I’m late to the thread, so ~nobody will see this. But it’s the culmination of about five yea…

FWIW, I see your comment. Also late to the thread though. This ruling is being watched at my office. I want to be a bit anonymous, but we've been doing a much more analogue version of some of these things for 75 years. With academics being our primary market. We've only had two legal issues in that time. Both settled out of court. But we walk a fine line.

Thank you for your work!

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#194
post #183

Earlier quoted context omitted.

Can you offer some examples from this ruling? It seems pretty reasonable on a first read.

Judge decided having an output filter on your AI makes it ok for it to contain full copy of copyrighted work. Its like saying it should be legal for me to have this Judges nudes obtained 100% illegally as long as I pixelate all the naughty bits.

Full ruling is here (https://storage.courtlistener.com/recap/gov.uscourts.cand.43...)

The analogy the judge gives is to how Google Books walked the tightrope on copyright: they maintain an archive of all the books for indexing and search purposes, and can display excerpts to help you confirm that's what you're looking for. The excerpts are constrained so you can't read the whole book by scanning the excerpts.

If post-filtering the LLM signal is illegal, shouldn't Google Books archive also be illegal? If not, why not?

And if you believe it should be, understand that the way precedent works, the judge won't be ruling that way without pulling some fire on themselves, because it is not the business of another case to contradict the conclusions of a previous court in a previous case. Copyright law is arbitrary and highly path-dependent because the underlying goal is forever in tension with itself, that goal being providing societal benefit by creating artificial scarcity on something that is, by its nature, not scarce at all.

(Worth noting: Anthropic didn't get off scot-free. The ruling was that the created artifact, the LLM, was a fair-use product, but the way it was created was through massive piracy and Anthropic is liable for that copying).

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#195

One aspect of this ruling [1] that I find concerning: on pages 7 and 11-12, it concedes that the LLM does substantially "memorize" copyrighted works, but rules that this doesn't violate the author's copyright because Anthropic has server-side filtering to avoid reproducing memorized text. (Alsup compares this to Google Books, which has server-side searchable full-text copies of copyrighted books, but only allows user…

A judge already ruled that models themselves don't constitute copyright infringement in Kadrey v. Meta Platforms, Inc. ( https://casetext.com/case/kadrey-v-meta-platforms-inc ). The EFF has a good summary about it: > the court dismissed “nonsensical” claims that Meta’s LLaMA models are themselves infringing derivative works. See: https://www.eff.org/deeplinks/2025/02/copyright-and-ai-cases...

Time to overfit on some books and publicize them as a libgen mirror.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#196
post #188

Earlier quoted context omitted.

I guess the optimistic take would be that we will get novel, insightful synthesis of disparate fields of knowledge that no human so far was ever able to hold in their mind to contemplate their interrelations. And this will elevate the human spirit etc. The equivalent to the take that the Internet will bring peoples together and foster better understanding and love between people who so far were not in dialogue and th…

I suppose another strongman would be that LLMs substantially decrease the cost of human creation (i.e. the HITL assistant use case) while producing an output of equivalent quality. As a result of this, everything gets cheaper and more plentiful. The counterargument I'd make to that would be the requirement that the human have creative skills, which might atrophy in the absence of business models supporting a career c…

Generally, having cheap mass produced things can be great compared to only expensive artisanal stuff that only the rich can afford. Think about furniture, clothes etc. or all the other stuff you have in the house, compared to 100-150 years ago. Today we can buy pretty good mass produced furniture for example. A few generations ago people either did it themselves in a wonky way or paid a lot of money for a hand made carpentry option. Just like with LLMs. LLMs probably do a better job in general writing than a random person off the street. But it's not as good as the top performers. But it's much cheaper. It's a tradeoff.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#197
post #167

Earlier quoted context omitted.

We spun "intellectual property" law from whole cloth. We'll need to reweave it now. Deal with it.

Rewriting does not mean destroying, as the cannibalization of news reporting by social media should have taught us. It's entirely possible for something to be suboptimal in the specific (I would like this thing for free), but optimal on the whole (society benefits from this thing not being free).

The societal benefits we've enjoyed from copyright law have been substantial, but the upside is completely maxed out at this point. The tail has been wagging the dog since the MPAA and RIAA grew into de-facto government agencies.

The potential societal benefits to AI are unbounded, but only if it's allowed to develop without restrictions that artificially favor legacy interests.

Any decision or legislation that says that training is not fair use -- and yes, that includes gaining access to the content in the first place by any means necessary -- will have net-negative effects on the society that enforces it.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#198
post #188

Earlier quoted context omitted.

I suppose another strongman would be that LLMs substantially decrease the cost of human creation (i.e. the HITL assistant use case) while producing an output of equivalent quality. As a result of this, everything gets cheaper and more plentiful. The counterargument I'd make to that would be the requirement that the human have creative skills, which might atrophy in the absence of business models supporting a career c…

Generally, having cheap mass produced things can be great compared to only expensive artisanal stuff that only the rich can afford. Think about furniture, clothes etc. or all the other stuff you have in the house, compared to 100-150 years ago. Today we can buy pretty good mass produced furniture for example. A few generations ago people either did it themselves in a wonky way or paid a lot of money for a hand made c…

The difficulty is the biggest gains there are for singular goods which can't be copied at low cost.

Exquisitely designed piece of furniture = expensive copy

Well-written book = cheap copy, post-printing press

So we're not necessarily going to get "more access to better" (because we already had that), but just "cheaper".

Whether that hollows out entire markets or only cannibalizes the bottom of the market (low quality/cheap) remains to be seen.

I wouldn't want to be writing pulp/romance novels these days...

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#199
post #167

Earlier quoted context omitted.

Rewriting does not mean destroying, as the cannibalization of news reporting by social media should have taught us. It's entirely possible for something to be suboptimal in the specific (I would like this thing for free), but optimal on the whole (society benefits from this thing not being free).

The societal benefits we've enjoyed from copyright law have been substantial, but the upside is completely maxed out at this point. The tail has been wagging the dog since the MPAA and RIAA grew into de-facto government agencies. The potential societal benefits to AI are unbounded , but only if it's allowed to develop without restrictions that artificially favor legacy interests. Any decision or legislation that says…

> The potential societal benefits to AI are unbounded, but only if it's allowed to develop without restrictions that artificially favor legacy interests.

That's a very strong claim based on currently limited evidence.

It's in no way clear that AI has an infinite ability to scale capability, nor that that can only be done by completely ignoring compensating those who provide training data.

OpenAI and Anthropic would love that to be true... but the facts don't support it.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#200

Earlier quoted context omitted.

[flagged]

We spun "intellectual property" law from whole cloth. We'll need to reweave it now. Deal with it.

The right to own property and the fruits of one's labor is a fundamental natural right, not something we "spun from whole cloth."

Deal with it.

Post reply on HN