Live data from Hacker News

AI companies are shredding rare books

twitter.com

21–30 of 559 posts

Re: AI companies are shredding rare books

#21
> A federal judge ruled the practice is fair use because eliminating the original means only one copy exists at a time.

Really? That sure wasn't a thing when one startup got sued for streaming from its wall od dvd-players, and they adhered to 1 disc = maximum 1 stream at same time.

Re: AI companies are shredding rare books

#22
post #12
post #4

Which book that was rare was destroyed? I'm interested to know a few titles.

The 404media article mentions notably https://nltimes.nl/2026/06/25/rare-book-dealers-fear-tech-fi... which says: > The attachment contained 3,000 English-language titles organized by ISBN number, including books such as Distinct Element Modelling in Geomechanics by K.R. Saxena (1999); Barrett's Traditional Fairy Tales (2021), an academic study of Irish folklore; and Laser Shock Peening of Advanced Ceramics by Pratik…

> Barrett's Traditional Fairy Tales (2021)

How is a book from 2021 considered rare in this context? There's almost certainly a digital copy of it in existence prior to Anthropic purchasing a print edition.

Re: AI companies are shredding rare books

#23

> A federal judge ruled the practice is fair use because eliminating the original means only one copy exists at a time. Really? That sure wasn't a thing when one startup got sued for streaming from its wall od dvd-players, and they adhered to 1 disc = maximum 1 stream at same time.

That's not an accurate characterization of the ruling

Re: AI companies are shredding rare books

#24
post #8

> You can reprint a bestseller. You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal. So it's going to accelerate. Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright? IDK about the shredding, it's not nice, but it's more a problem with copyright…

> Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright?

It's cheaper to scan the books if you do it destructively. Cost. That's why they're shredding irreplaceable texts. Nothing to do with copyright.

https://www.404media.co/ai-companies-are-buying-tons-of-old-...

Re: AI companies are shredding rare books

#25
post #20

I think people are, on the whole, too precious about old things. In the case of books produced after major commercial printing began, I don't believe it is the paper that imbues the book with historical value. Indeed, I think there's a high chance that this process increases preservation of the most relevant part of the media - the actual content! There are tons of old rolls of film slowly rotting away in warehouses…

So if the Mona Lisa is part of a model we can burn it?

If you purchased the Mona Lisa (or some rare book), in this hypothetical, you can burn it.

Re: AI companies are shredding rare books

#28
post #8

> You can reprint a bestseller. You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal. So it's going to accelerate. Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright? IDK about the shredding, it's not nice, but it's more a problem with copyright…

Scanning books by taking them apart into singular pages and scanning those pages is faster and cheaper. AI training is a numbers game, so they want faster and cheaper.

What happens to the pages after? No one needs them anymore, so they get mulched and recycled.

That would be the dominant scanning method even if copyright wasn't a thing. But then again - if copyright wasn't a thing, there would be much less need to scan any physical media.

The reason why OpenAI can't just go on Amazon, buy a "digital edition" of a 2018 book and use that is that it would violate the license in ten ways, and then the DMCA laws that forbid breaking DRM on top of it.

Re: AI companies are shredding rare books

#29

The publishers sued AI companies for training on shadow library data, hoping to negotiate content deals for big $$$ down the line. Instead, they got analog hole'd. Turns out that buying an old book for $5 and destructively scanning it for $25 is way cheaper than paying extortion fees to the copyright-mongers. What I don't buy is it being "rare, precious books". First, they're not after ancient texts - they're after t…

I see.. so the various AI companies are in the right on this?

It can be the case that everyone in “a fight” is wrong.

Re: AI companies are shredding rare books

#30
post #16

That's why archive.org should have never been sued for lending books they had physical copy of. This is the result. Publishers should be more careful what they wish for.

Publishers don't care if rare books get shredded?

Yeah, why would it be bad for publishers? If anything they'd most likely encourage more book shredding!
Post reply on HN