Live data from Hacker News

AI companies destroy physical books – let's scan rare books before it's too late

annas-archive.gl

191–200 of 961 posts

Re: AI companies destroy physical books – let's scan rare books before it's too late

#191

Getting 10 million people to do anything is really, really hard. Getting 10 million people to spend hours scanning a book (which takes a really long time with a home scanner) sounds impossible :(

Not everyone should scan stuff. If you spend any serious time looking through stuff that randos on the Internet have scanned the quality fits the Bell Curve perfectly.

Biggest problems:

  - scanning items that are bigger than the scanner platten so the start/end of every line is cut off.
  - becoming an "editor": scanning only the pages you think are interesting and skipping intros, forewords, title pages, copyright pages etc
I work in this space. I now require that before scanning a video is made carefully flicking through every page of the item so it can be checked after scanning to ensure all the pages are present and in the original order.

Even the big libraries fuck up. I wanted an intact copy of Harper's Weekly from 1900 that has a big fold-out map in it. It's not clear to the libraries scanning this issue that the map is missing from their copies. None of the copies for sale from dealers have the map. Even when it is still glued into the middle it gets missed by industrial scanners. Google's scan only includes the (blank) back of the folded map.

Luckily GPT was able to track down a copy in a university special collections and fired off an email asking them to scan it. I just got the scan today:

https://imgur.com/a/vgMkM7b

(preview size, they sent a 500MB TIFF)

Now I can reassemble the issue and upload it.

I spend a lot of tokens getting LLMs vision tools to find the missing pages in vintage items and then try to reassemble them from other scans where available.

I'm also splitting up volumes to reupload. A lot of periodicals are only available online as giant multi-gig volume PDFs with all the issues in one file. I have a separate app I wrote to scan all the pages looking for covers so they can be split into PDFs and then identifying the volume/issue/month/year data from the cover or title page.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#192
post #180

It would be nice if we have some "tracking" e.g. 30% of all known books are scanned. So far all information I searched in the Internet about the progress has been patchy. It's also impossible to understand if exact book was ever digitized or not.

Anna's archive estimates that they have preserved 16% of the world's books.

https://annas-archive.gl/faq

Re: AI companies destroy physical books – let's scan rare books before it's too late

#193

Google Books was a great resource until the lawyers got involved. I was able to find and download (one screenshot at a time) a rare family history. The author died 100 years ago. The published disappeared 80 years ago. But now Google has locked it behind a limited preview. Google probably has the best collection of high quality scans, followed by the Hathi Trust. None of which are useable by anyone outside of those s…

I agree they're the best currently available, but a lot of their scans are straight garbage and need to be redone.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#194

I'm interested in learning/teaching technologies. Naturally science fiction examples are interesting. I have found that often LLMs are familiar with the contents of SF books. But I feel poorer, almost deprived, by the fact that all the LLMs I've checked with have NOT been trained on the contents of Eon by Greg Bear.

Does it exist in the web at all? Could be the contents have not been scraped for the dataset yet.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#195

Google Books was a great resource until the lawyers got involved. I was able to find and download (one screenshot at a time) a rare family history. The author died 100 years ago. The published disappeared 80 years ago. But now Google has locked it behind a limited preview. Google probably has the best collection of high quality scans, followed by the Hathi Trust. None of which are useable by anyone outside of those s…

Does Anna's Archive have it now? They scrape tons of sources, Google Books and HathiTrust included.

A lot of Hathi is locked behind university and library access restrictions. I sometimes have to track down students or someone who has a local library card to get items I need.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#196
post #84

I am baffled at these practices and somewhere confused on what's the end game here? monopoly on information? altering data? exclusive subscription based knowledge? Feels like we have welcomed the AI era with open hands hoping( at-least assuming) that data democracy will be there, yet feels like its a long road!

They likely would not destroy the original books after scanning but apparently it's the legal way to do things because of copyright.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#197
post #114

Earlier quoted context omitted.

So, have you tried finding out what the programming was in October 1994? Or what cultural ephemera appeared in the TV guides of that era alongside the schedules? Either there's a copy for the week you want in an archive, or somebody's got one for sale, or most often neither. This can piss you off, if as it happened you had a reason to care.

To play devils advocate, completely on the terms of your argument, would it be better for that particular human artifact to be shredded and its contents melted into an anonymized data pool, or for it to exist in a museum archive, in its original form, such that future generations can better understand what it was like to be alive in 1994? I’d personally choose the latter, especially given that the 1994 tv guide is no…

A museum - or any other building - can only hold so much physical stuff. How much of it do you really want preserved? How do you choose what is preserved (it's an eventually inevitable choice)? Do you save the 1980s stuff but not the 90s? Or save every even/odd year? Some other method? How much direct access do you think people need to pre-digital history?

Re: AI companies destroy physical books – let's scan rare books before it's too late

#198
post #189

Earlier quoted context omitted.

Shouldn't they be allowed to do whatever they want with the copy they bought, and so legally own? Isn't that the entire purpose behind "copy right"?

They are, they are in no risk of being arrested. But that does not make it free of moral judgement.

Why not?

Re: AI companies destroy physical books – let's scan rare books before it's too late

#199
post #6

Earlier quoted context omitted.

It is also the case that the copyright holders are often putting restrictions around use of electronic forms that are driving the desire to use physical copies. I doubt AI companies would use a single physical book if they could avoid it - absent the legal cloud over electronic rights. I have no evidence but I can't help suspecting in part the publicity around this is driven in part by rights holders that want to for…

Ai companies don't use ebooks, because they are more expensive than second hand books

On Amazon right now, retail prices for e-book copies are higher than for the corresponding paperbacks.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#200

Earlier quoted context omitted.

But that scan is never made available to us in its original form. So it getting scanned by the AI company does nothing to preserve the book.

...because it is illegal to copy copyrighted material. 70 years later they might do it.

The thing I find most hypocritical though is that they are probably never share their libraries with anyone. After scanning, downloading, stealing, overloading websites and everything in between, to acquire enough data for their stupid machine, they're not going to share their data? I get that most of it can't be shared, but a lot can. There's no reason why you need to destroy multiple copies of a book from 1880, when it's free to share.

At the same time I can understand keeping track of when each books enters public domain might also be an absolute nightmare, and I wouldn't blame the AI companies for not wanting to deal with that. For the stuff they absolutely know is clear, they should provide dumps for everyone to download.

Post reply on HN