Live data from Hacker News

AI companies destroy physical books – let's scan rare books before it's too late

annas-archive.gl

541–550 of 961 posts

Re: AI companies destroy physical books – let's scan rare books before it's too late

#541
post #496

Earlier quoted context omitted.

>From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge. How tho? Ever sunday at the flea market, I see thousands of books that are rotting, hoping for someone to buy them or at least take t…

Those books at the flea market aren't "rotting", they are available for sale. You seem perfectly fine living in a world where your flea market is devoid of books.

A lot of them are rotting. We are not talking "one of the three living copies of the first edition of Joyce's Ulysses". Rather "1956 statistics of the cultive of yuca in 'some small village from Mexico': a boring analysis". Those books have value to train LLMs as they are 100% free of AI text, but has been collecting dust (or rotting) in someone's room for decades, and no human is buying them even for 10 cents.

Also, Anna's text implies that the books are scanned and then mischievously destroyed so nobody has access again to the content. That's not the case: the books are "destroyed" before scanning, by dissasembling them in pages so they can be feed to the scanner. Scanning while keeping the book intact is difficult, as you need to software-unwarp the page before OCR'ing it, and expensive as you either need specialized scanners or humans doing it.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#542
post #479

Earlier quoted context omitted.

Yep, this feels pretty much it. Looking at the "rare books", it was books that nobody would care about or would just rot away anyway. (The example book of Old books of agriculture is probably not that important today) Liberians have to accept that much of their job is sending books to be burned. A lot of them try their best to get people to be interested in older books, but they have to make way for "newer" books ins…

A quick search suggests that most "weeded" books are sold on or donated rather than burnt or sent to landfill. Do you have sources that say otherwise?

Just go to a local library and ask a librarian or anyone who has a lot of books and tries to give them away; sadly, in most cases, they pick the valuable ones, and the rest just get sent for destruction (Burning).

It is pretty standard procedure.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#543
post #217

Earlier quoted context omitted.

The thing I find most hypocritical though is that they are probably never share their libraries with anyone. After scanning, downloading, stealing, overloading websites and everything in between, to acquire enough data for their stupid machine, they're not going to share their data? I get that most of it can't be shared, but a lot can. There's no reason why you need to destroy multiple copies of a book from 1880, whe…

Since they seem to leapfrog each others’ models every few months, the training data is one of the few ways they can build competitive advantage, and that explains why they don’t share, even if we don’t have to like this.

Would it even be legal to share? I don't think it would be.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#544

Earlier quoted context omitted.

>From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge. How tho? Ever sunday at the flea market, I see thousands of books that are rotting, hoping for someone to buy them or at least take t…

> This way, it will live forever instead, scanned and trained on, conjoined with the rest of our knowledge in a magic machine. But that's just it, it won't live on forever because, from the perspective of preservation, training is a lossy, noninvertible transformation. The LLM cannot legally produce the book verbatim, it will only spit out a regurgitation of the information, chopped and mingled into a broad informati…

Are you under the impression they scan the book, train on it, then destroy the digital copy? Because that's not what's happening. They scan it, and hold it forever to train future models on. The scan still exists, not available to the general public but that's no different than if they had bought the books and kept them in a private library closed to the public.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#545
post #215

Earlier quoted context omitted.

Unfortunately, that's not the case in the United States. The LOC only selects around half of published books to be permanently held. The rest are disposed of (usually returning them to the publisher, donating them to a library, or destroying them).

They should send them to The Internet Archive instead.

Why aren't we storming the Library of Congress? They are the real villains here.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#546
post #299

I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collective responsibility. Thus, the preservation of this knowledge, its availability, and accessibility are t…

Surely they have the high quality scans, but there would probably be the same legal restrictions to just share the archive.

I'm only one person, but I scan old books that had an impact on me growing up, and upload them to archive.org. Thankfully there are others that do the same. (And to be sure, FWIW, these are books that have not been printed for about 50 years—I suppose the software community would call them abandonware.)

Re: AI companies destroy physical books – let's scan rare books before it's too late

#547

Earlier quoted context omitted.

Yep, this feels pretty much it. Looking at the "rare books", it was books that nobody would care about or would just rot away anyway. (The example book of Old books of agriculture is probably not that important today) Liberians have to accept that much of their job is sending books to be burned. A lot of them try their best to get people to be interested in older books, but they have to make way for "newer" books ins…

Well if you're appealing to librarians: "Hey, guys, when the librarians get pissed about the destruction of books, it’s time to put those listening ears on. Because we are very comfortable with the idea that books are tools that can be retired. What’s happening right now is not that...." https://bsky.app/profile/annabookwriter.bsky.social/post/3mt...

[deleted]

Re: AI companies destroy physical books – let's scan rare books before it's too late

#548
post #299

I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collective responsibility. Thus, the preservation of this knowledge, its availability, and accessibility are t…

From their perspective, it's training data that their competitors don't have. If they make it available, they fill in their moat.

I love how sci-fi authors like Ray Bradbury toyed around with a similar issue but then got it so wildly wrong.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#549
post #299

I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collective responsibility. Thus, the preservation of this knowledge, its availability, and accessibility are t…

A small change in the copyright law would fix this problem. Something like: If a company is scanning material protected by copyright, it has to send a digital copy of the scanned material to Library of Congress within 5 working days.

The Library of Congress already has a copy of every book published in the US. How would this help?

Re: AI companies destroy physical books – let's scan rare books before it's too late

#550

Earlier quoted context omitted.

A small change in the copyright law would fix this problem. Something like: If a company is scanning material protected by copyright, it has to send a digital copy of the scanned material to Library of Congress within 5 working days.

Wait, so the library of congress is suddenly responsible for probably petabytes a day of incoming scans? To what end? Do they have to index it and make it available? Do they have to check the accuracy and integrity of the scans? How does this help anything, except create more work to throw in the trash?

https://www.google.com/search?q=how+many+books+are+published... https://www.google.com/search?q=how+many+megabytes+average+b... https://www.google.com/search?q=what+is+3+megabytes+times+2....
Post reply on HN