Live data from Hacker News

AI companies are shredding rare books

twitter.com

221–230 of 559 posts

Re: AI companies are shredding rare books

#221

Earlier quoted context omitted.

Niche text. It's not impossible that there was only ever under a thousand of them printed and released into circulation. A digital copy would exist somewhere, of course. But for us, that only matters if we can buy or download it. And for AI companies, that only matters if they can get a digital copy DRM-free and licensed permissively enough.

This argument doesn't make any sense. All manner of AI companies just ingest whatever random text they can find on the internet to train their data, including copyrighted publications. Why would DRM on a digital copy of a book matter?

> Why would DRM on a digital copy of a book matter?

Because DRM is just a way to make "breaking copyright" more practically cumbersome. What's easier, breaking digital DRM for each and every E-book you find, or just establishing a single pipeline for scanning physical books?

Re: AI companies are shredding rare books

#222

The design of copyright has always been to restrict the dissemination of knowledge so that somebody can turn a profit. The handwaved justification is that the profit encourages the creation of new knowledge whose dissemination can be restricted, but that still doesn't erase the fundamental dynamic. This is merely the latest incarnation. We can imagine a slightly different process on a few fronts - AI companies pay to…

This is the correct take.

Copyright has done more than anything else to prevent preservation and dissemination of knowledge. And it's forcing Anthropic's hand now. Although they could take more effort to preserve the books.

Re: AI companies are shredding rare books

#223

Earlier quoted context omitted.

That is literally what this article is about.

The "article" in question is a tweet. The "rare botanical text" is just an example the author of the tweet made up in order to generate sympathy.

The tweet is about a 404 Media article about the practice. It's linked elsewhere here.

Re: AI companies are shredding rare books

#224

Earlier quoted context omitted.

It’s been determined that training on lawfully acquired works is fair use. Presumably in this discussion of shredding physical books Dario and Sam are not pulling heists at the local library. I’m sure there’s ongoing litigation, and better sources than this, but fair use was determined in June 2025 in a sf federal district court https://www.goodwinlaw.com/en/insights/publications/2025/06/... Similar conclusion vs met…

[flagged]

If you're going to call people ignorant, you should probably know how to spell the name of the character you're referencing.

Re: AI companies are shredding rare books

#225

I've limited sympathy for the publishers. It pisses me off to reflect that they can sit on works until copyright expires, keeping them out of print. There's no real need for any of these so-called rare books to be rare while they're under copyright. And related to this, the books that are in print are mostly only in print in the shittiest way. I often see well-made books from the 17th or 18th centuries which are stil…

> I often see well-made books from the 17th or 18th centuries which are still in good nick At the risk of stating the obvious, any poorly made books from then wouldn’t have lasted this long and so you would never see them.

https://en.wikipedia.org/wiki/Survivorship_bias

Re: AI companies are shredding rare books

#226
post #204

Earlier quoted context omitted.

No, courts so far in the jurisdictions which have heard such cases, have ruled it's fair use. There is plenty of ongoing litigation in many jurisdictions, so it's way too early to just decree "it's been determined". It likely won't be for years to come.

Why did Anthropic settle with authors for 1.5 billion then? Surely their lawyers must have decided there's a pretty good chance of judges ultimately deciding that it is copyright infringement?

Anthropic settled because even though the training is fair use, Anthropic did not acquire all of the training material through legal means.

Re: AI companies are shredding rare books

#227
post #204

Earlier quoted context omitted.

No, courts so far in the jurisdictions which have heard such cases, have ruled it's fair use. There is plenty of ongoing litigation in many jurisdictions, so it's way too early to just decree "it's been determined". It likely won't be for years to come.

Why did Anthropic settle with authors for 1.5 billion then? Surely their lawyers must have decided there's a pretty good chance of judges ultimately deciding that it is copyright infringement?

They had illegally obtained books. 1.5 billion is an incredible deal compared to the per book infringement fine. https://fortune.com/2026/07/21/anthropic-copyright-settlemen...

Re: AI companies are shredding rare books

#228

Earlier quoted context omitted.

I have even less sympathy for IP stealing LLM operators

It’s been determined that training on lawfully acquired works is fair use. Presumably in this discussion of shredding physical books Dario and Sam are not pulling heists at the local library. I’m sure there’s ongoing litigation, and better sources than this, but fair use was determined in June 2025 in a sf federal district court https://www.goodwinlaw.com/en/insights/publications/2025/06/... Similar conclusion vs met…

Might surprise people to learn but the law isn't the final arbiter on what is moral or just. The law only decides what is legal and as we've seen over our lived history as humans, many "legal" things may not be those we want society to uphold.

Re: AI companies are shredding rare books

#229
It's not like these books were available to everyone before. If they are destroying one physical copy that's not accessible to the public and replacing it with a scanned copy that isn't accessible to the public, that's not a huge change. Except that will probably last longer digitized. Ideally they wouldn't destroy the physical copies, but this isn't anywhere as bad as book burning in Nazi Germany.

I find going after shadow libraries to be much worse because law enforcement is trying to prevent discrimination of knowledge to the public.

The real blame here should be going onto copyright laws.

Anthropic could take more care by figuring out if the books are still affected by copyright.

But this is just a company trying its best in an unfortunate regulatory environment.

Support your local shadow library:

https://annas-archive.pk/donate

Re: AI companies are shredding rare books

#230

They should be forced to publicly release the books as an Ebook. Leave it up for anyone to download and then compensate the copyright holders later. In fact if ingesting these books for LLMs is fair use, us commoners should be able to read them for free. Maybe restrict commercial redistribution though.

There's a bunch of details regarding LLMs and copyright, but I don't see how > They should be forced to publicly release the books as an Ebook. would be reasonable in any way? The books aren't theirs to release publicly. If I brought a copy of any given movie on DVD, ripped it and used it to train my own "LLM" located at /dev/null, should I then be allowed (or even forced) to release the movie publicly for anyone to…

We’re in a brave new world now. If theirs reason to believe your destroying the last copy of a book( or another copy isn’t easy to obtain) then you should make a copy of it available.

Set up a compensation fund for the rights holders. Anything is better than culture literally being sucked into the void.

Post reply on HN