Live data from Hacker News

AI companies are shredding rare books

twitter.com

131–140 of 559 posts

Re: AI companies are shredding rare books

#131

The publishers sued AI companies for training on shadow library data, hoping to negotiate content deals for big $$$ down the line. Instead, they got analog hole'd. Turns out that buying an old book for $5 and destructively scanning it for $25 is way cheaper than paying extortion fees to the copyright-mongers. What I don't buy is it being "rare, precious books". First, they're not after ancient texts - they're after t…

Buying up old books is legal. Digitizing and distilling them arguably is too. And also:

> paying extortion fees to the copyright-mongers

Yes yes, greedy fat cat publishing oligarchs treading on the poor put-upon scrappy AI underdogs. /s

Back in reality, the fraction of people who got into publishing books to get rich collecting rents is… not large. There’s so many other fields that are likely to reward participants with more wealth that it’s absurd — even with all the passion for the work in tech it’s probably relatively less pure.

And whatever the excesses of copyright have been, the whole bargain has always been on more pro-social foundations and stronger intellectual foundations than “extortion” sneers. It recognizes that incentives matter and work that’s valuable should be rewarded and incentivized.

A culture that takes a Robin Hood approach to low marginal cost billing points but fawns over the hypercapitalized distribution King Johns isn’t creating a freer or richer society or fighting the real cartel center, it’s indulging resentment and caricature.

Re: AI companies are shredding rare books

#132

Earlier quoted context omitted.

The archive.org story was more nuanced than that. If I recall correctly the full story was that they used to lend digital versions of books they physically bought and scanned with DRM to enforce a sort of one to one at a time restriction. But during covid archive.org decided to just remove the limit and lend unlimited copies concurrently which started the debacle with the publishers.

> But during covid archive.org decided to just remove the limit and lend unlimited copies concurrently IIRC, this was 100% it. Lending one digital version of one physical asset was likely already a violation copyright. Lending UNLIMITED digital versions of one physical copy was DEFINITELY a blatant violation of copyright.

Correct, it was switching to unlimited lending instead of one lend per physical book that got them in trouble.

Re: AI companies are shredding rare books

#133

The publishers sued AI companies for training on shadow library data, hoping to negotiate content deals for big $$$ down the line. Instead, they got analog hole'd. Turns out that buying an old book for $5 and destructively scanning it for $25 is way cheaper than paying extortion fees to the copyright-mongers. What I don't buy is it being "rare, precious books". First, they're not after ancient texts - they're after t…

Modern texts become ancient texts; given enough time.

Re: AI companies are shredding rare books

#134
People keep bringing up these rare books but never share what they actually are. What are their names? When were they written? How many copies were in existence? Were they in libraries, or locked up in vaults? Did people have access to them prior to being scanned for AI? How many such books have actually been destroyed?

Weird to see so many of these "trust me bro" twitter stories make it to the front page and cause outrage when no one has any real information.

Re: AI companies are shredding rare books

#135

We reprint old books after checking out copyrights (for all books, this means pre-1930, but for some (I'd actually say most) it also means ones published up to 1964 and 1973, depending on how the rightsholders did (or didn't) do the renewals). We use a special guillotine type cutter to cut off the binding and then store the pages in a sealed plastic bag which goes in the archives; they're stored there indefinitely in…

Why cut off the spines? Isn't that how you end up getting unattributed 'dead sea scrolls'?

It's easier to get good quality scans from individual pages than it is from pages in a complete book. Imagine laying a book flat, then the page is distorted in the area around the spine. You can work around this (either by trying to correct for the distortion in software or with clever scanners that position the book more advantageously) but it adds complexity compared to chopping off the spine and just dealing with flat sheets of paper.

Re: AI companies are shredding rare books

#139
post #12
post #4

Which book that was rare was destroyed? I'm interested to know a few titles.

The 404media article mentions notably https://nltimes.nl/2026/06/25/rare-book-dealers-fear-tech-fi... which says: > The attachment contained 3,000 English-language titles organized by ISBN number, including books such as Distinct Element Modelling in Geomechanics by K.R. Saxena (1999); Barrett's Traditional Fairy Tales (2021), an academic study of Irish folklore; and Laser Shock Peening of Advanced Ceramics by Pratik…

Non of those are rare. All a available in libraries for ILL.

Re: AI companies are shredding rare books

#140

Earlier quoted context omitted.

> But during covid archive.org decided to just remove the limit and lend unlimited copies concurrently IIRC, this was 100% it. Lending one digital version of one physical asset was likely already a violation copyright. Lending UNLIMITED digital versions of one physical copy was DEFINITELY a blatant violation of copyright.

How is lending one digital version of one physical asset a violation of copyright? Since I MAY be able to lend out the physical as well?

From the court decision:

"IA maintains that it delivers each Work “only to one already entitled to view [it]”―i.e., the one person who would be entitled to check out the physical copy of each Work. But this characterization confuses IA’s practices with traditional library lending of print books. IA does not perform the traditional functions of a library; it prepares derivatives of Publishers’ Works and delivers those derivatives to its users in full. That Section 108 allows libraries to make a small number of copies for preservation and replacement purposes does not mean that IA can prepare and distribute derivative works en masse and assert that it is simply performing the traditional functions of a library. 17 U.S.C. § 108; see also, e.g., ReDigi, 910 F.3d at 658 (“We are not free to disregard the terms of the statute merely because the entity performing an unauthorized reproduction makes efforts to nullify its consequences by the counterbalancing destruction of the preexisting phonorecords.”)."

Post reply on HN