Live data from Hacker News

The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions

dallasexpress.com

41–50 of 79 posts

Re: The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions

#41
I'm not really the most AI-friendly person around but this news cycle is just as equally annoying.

"Rare and out-of-print" is fabricating a lot of aura here. It's technically correct (the best kind of correct) but I've yet seen evidence that these are culturally significant copies being destroyed for scanning.

From https://nltimes.nl/2026/06/25/rare-book-dealers-fear-tech-fi...:

> The attachment contained 3,000 English-language titles organized by ISBN number, including books such as Distinct Element Modelling in Geomechanics by K.R. Saxena (1999); Barrett's Traditional Fairy Tales (2021), an academic study of Irish folklore; and Laser Shock Peening of Advanced Ceramics by Pratik Shukla (2018).

The oldest is from 1999! If that's the best they can actually enumerate to bolster this outrage farming cycle you could just wonder how irrelevant the rest are.

Really, please, kill this news cycle. There's a lot of issues deserving proper attention right now and this one is a straight-up nothingburger.

Re: The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions

#43
post #27

Earlier quoted context omitted.

One (1) company destroys and (illegally?) copies a rare book. This knowledge is now of that company, not of humanity. Other companies will feel the need to do the same and before you know it knowledge is walled off in the gardens of the AI companies, and the rare books are gone and destroyed.

The scan doesn't go away, and can be released when the copyright expires. Oh no, a lump of cellulose is gone. Will no one think of the fibers?

> can be released when the copyright expires

But it ... won't be? The companies have no motivation to do so. The copyright on lots of these books are surely already expired, so if they wanted to they could be putting these up now. I'm sure internally this is viewed as a corpus of knowledge they have that their competitors do not, so they will not release them unless something forces them to.

Re: The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions

#44

On the surface it seems bad, although many of these books were probably rotting in place rather than being read. If the companies are willing to make the digital version available this may actually prreserve the books. More troubling to me is the focus on pre-2022 data. This indicates that the companies are seriously worried about model collapse, where the future models become worse becuase they are being fed by gene…

[deleted]

Re: The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions

#45
post #34
post #31

Humanity hasn't caught up to the reality of the times; we should adjust copyright laws so anything over say 10 years old can be hosted for free on the Internet. We should have colossal databases of music, books, movies, etc. that can be freely traded and downloaded without breaking any laws. Maybe in the past the commerical protection of art and IP was necessary but with the Internet and AI I think we need to, as a s…

10 years may be too little but generally I agree. What's blowing my mind lately though is the people complaining about copyright today are some of the same people I saw 30 years ago screaming "Information wants to be free" and running around with the DeCSS code on their T-Shirts. But suddenly, now that they hate AI, copyright is important and the information should be locked away. I can't wait for this all to settle…

It's possible to be opposed to excessive copyright and also be opposed to the mass theft of media for profit and subsequent abuse of copyright that the scrapers are guilty of. The AI pillagers are taking everything in, then gatekeeping access. And as they acknowledge with their "nothing after 2022" preferences, they (or their users at least) are salting the earth for future seekers of knowledge.

Re: The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions

#47

There are plenty of scanners that do not destroy the book being scanned. Here is one, I am sure there are many others: https://www.youtube.com/watch?v=b9LcTZU-HHI

It should be more than obvious, that the problem are not the scanners.

Re: The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions

#48
post #45
post #34

Earlier quoted context omitted.

10 years may be too little but generally I agree. What's blowing my mind lately though is the people complaining about copyright today are some of the same people I saw 30 years ago screaming "Information wants to be free" and running around with the DeCSS code on their T-Shirts. But suddenly, now that they hate AI, copyright is important and the information should be locked away. I can't wait for this all to settle…

It's possible to be opposed to excessive copyright and also be opposed to the mass theft of media for profit and subsequent abuse of copyright that the scrapers are guilty of. The AI pillagers are taking everything in, then gatekeeping access. And as they acknowledge with their "nothing after 2022" preferences, they (or their users at least) are salting the earth for future seekers of knowledge.

> It's possible to be opposed to excessive copyright

Define excessive. Who makes that determination and how?

Re: The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions

#49
post #31

Humanity hasn't caught up to the reality of the times; we should adjust copyright laws so anything over say 10 years old can be hosted for free on the Internet. We should have colossal databases of music, books, movies, etc. that can be freely traded and downloaded without breaking any laws. Maybe in the past the commerical protection of art and IP was necessary but with the Internet and AI I think we need to, as a s…

I imagine you're in favor of the artists being compensated in some other way then? Or should they just suck it up?

Re: The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions

#50
post #34
post #31

Humanity hasn't caught up to the reality of the times; we should adjust copyright laws so anything over say 10 years old can be hosted for free on the Internet. We should have colossal databases of music, books, movies, etc. that can be freely traded and downloaded without breaking any laws. Maybe in the past the commerical protection of art and IP was necessary but with the Internet and AI I think we need to, as a s…

10 years may be too little but generally I agree. What's blowing my mind lately though is the people complaining about copyright today are some of the same people I saw 30 years ago screaming "Information wants to be free" and running around with the DeCSS code on their T-Shirts. But suddenly, now that they hate AI, copyright is important and the information should be locked away. I can't wait for this all to settle…

People must realize that it is easy to built some sentiment against on part of legislation that seems to restrict them, however, it is much more difficult to establish legislation that serves everyone. In the end we need smart people in politics and lawmaking instead of populists and libertarian cry-babies.
Post reply on HN