Live data from Hacker News

AI companies destroy physical books – let's scan rare books before it's too late

annas-archive.gl

761–770 of 961 posts

Re: AI companies destroy physical books – let's scan rare books before it's too late

#761

The scale of problem seems a bit overblown. Anna's Archive paint a picture like AI companies are some movie villains burning books so no one can see them. but in reality they just disassemble them into pages because it's cheaper and faster to scan. Most of these books is highly specialized, they been collecting dust on shelves for decades and nobody need them. But the problem is real. Even if these books aren't neede…

> Anna's Archive paint a picture like AI companies are some movie villains burning books so no one can see them. but in reality they just disassemble them into pages because it's cheaper and faster to scan. Most of these books is highly specialized, they been collecting dust on shelves for decades and nobody need them.

"Paint a picture" is apt.

"... like AI companies are some movie villains burning paintings so no one can see them. but in reality they just cut the canvas off the frames because it's cheaper and faster to photocopy. Most of these paintings have been collecting dust in galleries for decades and nobody needs them."

Re: AI companies destroy physical books – let's scan rare books before it's too late

#762

Earlier quoted context omitted.

If they're 50 years old they're young, and archive.org will likely block access. If they're not already on annas-archive (or the copy there is trash), your best bet is an anon upload to libgen.

Thanks. There was a time of course when you could pull my books down from archive.org as PDFs. Perhaps that time will come again. I'll look into libgen.

Archive.org Scanned Book Downloader Bookmarklet

https://gist.github.com/cemerson/043d3b455317d762bb1378aeac3...

Re: AI companies destroy physical books – let's scan rare books before it's too late

#763
post #733

Nondestructive scanning can cost 10x as much. This is about cost. It is not about preservation. Google never destroyed the books it scanned. Amazon and Anthropic are attempting to save money. They are not considering whether or not a book is rare. They are treating books as a commodity. Rare books are rare. It is easy enough to identify when there are a limited number of copies of a book. The issue is saving money on…

Given all the bad PR around this issue - you'd think they'd do the 10 seconds of examination "is this rare" before putting it in the cutter. I'm genuinely surprised they don't.

I'm surprised you're confident they don't.

You think someone's picking up a first edition Steinbeck and destroying it so it gets into the next training corpus?

Or is it like, technical manuals and incredibly dry almanac content?

I think the details actually matter here. I'd love to see a realtime list of these titles being scanned and processed. Then we'd know if anyone actually gives a care.

That said I'm willfully ignorant of most "headline news" so I'd be curious to see how heinous the issue really is.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#764
post #251

Aren't the AI companies just buying one copy of each book? And this is supposed to be concerning?

The blog post is about rare books. Meaning, AI companies scanning and destroying rare books.

Yes, read more carefully: one copy. Do you panic when the "wrong person" buys a single rare book?

The issue is that the copyright holders and book publishers make it hard to make more copies, not that somebody bought a single copy of a book, no matter how rare.

If you could print any book on demand (paying for it), the issue would vanish instantly. And who makes that decision?

Re: AI companies destroy physical books – let's scan rare books before it's too late

#765

You ask "Why destroy physical books?" I ask "Why save physical books?" If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere. Owning and storing physical books is not free, there is a real cost. Even owning and storing the scans is not free, especially when IP rights and challenges get involved, since a scanned book with no distribut…

> If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.

That is literally the opposite of how it works.

https://en.wikipedia.org/wiki/Scarcity

Re: AI companies destroy physical books – let's scan rare books before it's too late

#766
AI companies are supporting the market for books that no one else wanted. The economically illiterate assumption of the author is that "rare" books are good. If they were that good, they would be priced higher.

https://abunner.substack.com/i/210907372/anthropic-is-suppor...

Re: AI companies destroy physical books – let's scan rare books before it's too late

#767

I don't see any mention of Project Ocean - AKA Google books. Before AI they endevoured to digitize books in a massive online library. This inccluded rare and out of print books many which are archived at libraries. Because they had to preserve the books and return them in the condition they received then they created elaborate technology to accomplish this. The project was met with significant legal challenges from a…

Whenever I need something from Google Books I inevitably reach the message that this is a limited preview and the part I need is not included. I therefore feel the same way about Google Books that how I felt when I learned that What.cd went down: that I don't gain or lose anything anyway because I never had access to begin with, and that by not making it 100% publicly accessible you're asking for the data to one day…

> Whenever I need something from Google Books I inevitably reach the message that this is a limited preview and the part I need is not included.

> I therefore feel the same way about Google Books that how I felt when I learned that What.cd went down: that I don't gain or lose anything anyway because I never had access to begin with, and that by not making it 100% publicly accessible you're asking for the data to one day disappear forever.

Can you still search the restricted parts? If so there's still value to it: it helps you identify the book so you do an inter-library loan to get at the full content. Sure, it's not frictionless, but I wouldn't be all or nothing about it.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#768

Earlier quoted context omitted.

bro if people on HN of all places are concerned of physical media being destroyed boggles ur mind then u need to read more about the damage caused to society by prior happenings. I wasn't even making an unusual or particularly strong statement Im going to try and be patient and explain my position. I just feel that actions that are destructive should always be seen through a lens of suspicion. IE, lets destroy this w…

Nothing is destroyed, only transformed. What was a physical book is now a digital book. It's a completely different story from "prior happenings", to which I assume you mean book burning and the like. Pretending they are the same is intellectually lazy.

> Nothing is destroyed, only transformed. What was a physical book is now a digital book.

I genuinely don't understand how it's possible for someone to type that out in all sincerity. Have you ever held a physical book?

Re: AI companies destroy physical books – let's scan rare books before it's too late

#769

Big AI companies are leaving an easy opportunity on the table for establishing goodwill with the public. Just publicize a rare books vault where you put the older editions that aren’t in a lot of library catalogs. Use non-destructive scanning for those. Align yourself with the image of safeguarding something. It seems like a no-brainer given various themes I’ve been hearing in criticisms of these companies. Maybe the…

From what I understand, to work with copyrighted books they need to essentially format shift (i.e., scan and destroy the physical book). So a book vault would not solve this issue. A book vault would still be useful for out-of-copyright works, but this would only cover a (probably relatively small) portion. Also, I'm not sure how easy it is to reliably determine copyright at scale, so they might just decide that it's…

They could put everything in some kind of nonprofit book vault/archive that includes all the source pages as long as they didn't use or distribute it. They could even provide a mechanism for rights holders to recover the text for free, if they need access.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#770

People keep repeating the "rare books" without providing any evidence that they are rare. Anyone who has collected books knows there are massive volumes of old books that can be bought by the pound.

Any named book I have seen in the coverage of this are cheap crap books, that are available somewhere in the word from a public library. Its literal trash.
Post reply on HN