Live data from Hacker News

AI companies destroy physical books – let's scan rare books before it's too late

annas-archive.gl

581–590 of 961 posts

Re: AI companies destroy physical books – let's scan rare books before it's too late

#581
Why are these books "rare"? Because no one wants them. Why then are we up in arms over their destruction? These sensational headlines make it seem as though a copy of the Codex Sassoon 1053 is being destroyed, when in fact these are just obscure books that no one cares about.

The rhetoric on this topic is reminiscent of the rhetoric regarding data centers: some noxious combination of misinformation, misunderstanding, and sensationalism, wielded against technological progress.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#582

Earlier quoted context omitted.

A small change in the copyright law would fix this problem. Something like: If a company is scanning material protected by copyright, it has to send a digital copy of the scanned material to Library of Congress within 5 working days.

Puts the burden on government to store what is probably 90% worthless material. Copyright should really be amended so that once out of print and a grace period it’s free use. I am probably more of an anarchist in this regard. Similar to my belief that anyone should be able to ingest any data you put online, once a book is no longer being print it should be able to be used for commercial or personal use for free. Simi…

> Puts the burden on government to store what is probably 90% worthless material.

That's what governments are for.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#583

Earlier quoted context omitted.

I can't tell if you're making a joke but many (most?) rare books predate the modern copyright regime and the original printing plates are somewhere in a 17th century midden heap.

The books being destroyed are not afaik those sort of books. They are just out of print modern books, that are still under copyright (and hence can't legally be scanned non-destructively).

> They are just out of print modern books, that are still under copyright

OR in-print modern books that they can get for cheaper by buying used. The whole thing is a manufactured outrage over something that doesn't matter. They aren't breaking into museums to steal their only copy of a book and burn it. They are digitizing books. If anything I commend them for what they are doing. If the alternatives were that book rotting on a shelf or being thrown away they doing a great service preserving it, even if they don't make it available publically (which they can't for copyright reasons). It's literally no different from them stocking a private library with these books, except it's better because digital copies are much easier to preserve.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#585

I don't see any mention of Project Ocean - AKA Google books. Before AI they endevoured to digitize books in a massive online library. This inccluded rare and out of print books many which are archived at libraries. Because they had to preserve the books and return them in the condition they received then they created elaborate technology to accomplish this. The project was met with significant legal challenges from a…

Whenever I need something from Google Books I inevitably reach the message that this is a limited preview and the part I need is not included.

I therefore feel the same way about Google Books that how I felt when I learned that What.cd went down: that I don't gain or lose anything anyway because I never had access to begin with, and that by not making it 100% publicly accessible you're asking for the data to one day disappear forever.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#586

Earlier quoted context omitted.

Surely they have the high quality scans, but there would probably be the same legal restrictions to just share the archive.

I'm only one person, but I scan old books that had an impact on me growing up, and upload them to archive.org. Thankfully there are others that do the same. (And to be sure, FWIW, these are books that have not been printed for about 50 years—I suppose the software community would call them abandonware.)

If they're 50 years old they're young, and archive.org will likely block access. If they're not already on annas-archive (or the copy there is trash), your best bet is an anon upload to libgen.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#587
What’s really disgusting is how unnecessary this is.

LLMs have topped out in terms of language fluency. You’re not going to get a smarter model with 250 trillion tokens than with 25 trillion tokens. There are still other gains to be made in the LLM/LRM space, but they don’t require ripping up rare books.

And they’re doing it destructively because it’s cheaper. That’s it. They absolutely could scan nondestructively. They’re trillion-dollar companies, and they do this in a shitty way to save pennies.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#588
post #120
post #77

Earlier quoted context omitted.

That would make sense as an argument if the natural endpoint of these copies was preservation, and AI was disrupting that. But the natural next step for virtually all these books is to be recycled, not preserved. Books are generally not preserved . It is extremely normal for them to be pulped. Millions and millions of books are pulped every year.

Might as well grind up the tablet of complaint to Ea-nāṣir and make cement out it, right? What use could there be in preserving the mundane facets of everyday existence?

When applied to ordinary mundane objects, this is the mindset that leads to pathological hording behavior. Ultimately I think it's rooted in a fear of, an attempt to deny, mortality and the passage of time.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#589
post #407

How is Anna's Archive getting around the copyright violations of hosting all these books for access to all? I suspect it won't be long before they get sued and are forced to shut down. I spent some time reading the web site, and it doesn't look to be a well thought out project. Even the way it is organized leaved much to be desired. There's much more to library science and the organization of a vast collection of boo…

Unlike archive.org which is a real business, Anna's Archive is basically piracy. You could try to sue them but you'd have to find them first. This IMO makes them more resilient against that sort of thing, and also means they can actually do their job.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#590
post #78

Earlier quoted context omitted.

I am thinking about a long essay Icelandic author Þórbergur Þórðarson wrote to his pals abroad, and were left abroad. I am thinking about a photo book by an Indonesian naturalist who is famous on Bali, and took amazing photos of wildlife on Sulawesi in 1926, and colored in, and somehow ended up in New Jersey in the 1980s. I am thinking about a collection of essays written by a teenage J.D. Salinger who he left unsign…

You're trying to imagine rare valuable books and then fantasizing about AI companies destroying them, but what's actually happening here is that AI companies are acquiring, digitizing, and then pulping the instruction manuals to 1983-vintage copy machines. This is all such a special-pleading argument. You know what other institution snatches up books and destroys them at huge scale? Public library systems. People cle…

I assume public libraries know what they are doing and know which book they are destroying, that they keep an accurate inventory and hire professionals maintaining said inventory and marking which books can safely be destroyed.

I assume no such things of AI companies.

Post reply on HN