Live data from Hacker News

AI companies destroy physical books – let's scan rare books before it's too late

annas-archive.gl

651–660 of 961 posts

Re: AI companies destroy physical books – let's scan rare books before it's too late

#651

Earlier quoted context omitted.

Nothing is destroyed, only transformed. What was a physical book is now a digital book. It's a completely different story from "prior happenings", to which I assume you mean book burning and the like. Pretending they are the same is intellectually lazy.

Thank you for your reply and request for clarification. Yes destroying a book is in my opinion identical to a book burning. One day there will be no hard copies, the digital books are hence vulnerable to patches, whether its due to a new political movement or a sudo abled hamster running on a keyboard... and if that occurs information will be permanently lost. I would have thought a HN user understand well the import…

There are multiple digital ways to preserve exact copies of the original. Both through things like personal backups (I have digital backups of all my physical books) and through collective recordkeeping. For example, Anna's Archive stores them (or at least provides access) via a hash of the book. You can't patch the book and keep the same hash. Yes, hash collision is a thing but there are better hashes or other ways to accomplish the same goal. It pains me to say "blockchain" but that's something we have today that could be leveraged, though I think there are other cryptographic methods we could employ.

I'll admit I have less concern with (every) digital copy being altered and if we assume a future where all computing is completely locked down and governments have the ability to reach in a tweak anything then yes, we are screwed. But we are screwed if we allow that to happen even if some of us have squirreled away physical books. In that techno-hell (of completely locked down computing with no open options) then some of use will have "illegal" computers that still do what we want and I see that as no different than keeping physical copies.

Put simply: Preserving the original does not require a physical substrate (or at least one made of paper, obviously computers run on physical hardware).

Re: AI companies destroy physical books – let's scan rare books before it's too late

#652

Earlier quoted context omitted.

That sounds to me like fuzzily having a version of the book. That is, it's more like having the book, than having nothing would be. You seem to be arguing that a translation, etc is not literally having the book, which is something that has always been my stance as well.

You're trying to use the Socratic method to show that the havingness of the book is a spectrum, and I'm taking the position of a hardliner who considers that "having the book" means having the original string of symbols from beginning to end, and anything besides that is not having the book. I'm trying to show you that your line of argumentation is uncompelling to such a person.

I was never using the Socratic method. I was asking rhetorical questions, and then gave the (my) answer at the end.

We just have different opinions about what it means to have a book. I think that having a copy of "Pride and Prejudice and Zombies" is more like having a copy of "Pride and Prejudice", than having no book at all is like having a copy of that book. You disagree and that's fine.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#653

Earlier quoted context omitted.

Can you explain why AI companies destroying these is worse than a library or bookseller destroying these?

Public libraries staff professional archivist, they take careful inventory and know exactly which book they are destroying. AI companies buy boxes and boxes of books with unknown content staff anybody who can operate a scanner and have no idea which books they are destroying. If a public library comes across a book they didn’t know they had, it is very likely that somebody will notice and know how to continue, to fin…

At this point I believe you're just sort of making up a fantasy of what you think a public library should do, not what they actually do.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#654
The scale of problem seems a bit overblown. Anna's Archive paint a picture like AI companies are some movie villains burning books so no one can see them. but in reality they just disassemble them into pages because it's cheaper and faster to scan. Most of these books is highly specialized, they been collecting dust on shelves for decades and nobody need them.

But the problem is real. Even if these books aren't needed by anyone right now, them digitization in a single copy that end up behind seven locks at a corporation is not great, because AI doesn't replace the original. You can't to ask a neural net to give you a exact copy of a page from that book. So yeah, the post dramatizes a bit, but the point are valid. We need open digital archives.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#655

Earlier quoted context omitted.

Whenever I need something from Google Books I inevitably reach the message that this is a limited preview and the part I need is not included. I therefore feel the same way about Google Books that how I felt when I learned that What.cd went down: that I don't gain or lose anything anyway because I never had access to begin with, and that by not making it 100% publicly accessible you're asking for the data to one day…

Quod licet Iovi, non licet bovi Big companies will read up the books and make their AI recite them from memory, but Archive.org was sued for renting one book on an exclusive basis (unless one would return, another wouldn't be able to rent)

> Archive.org was sued for renting one book on an exclusive basis (unless one would return, another wouldn't be able to rent)

No, this is what they were doing before, but they explicitly started lending out "unlimited" copies, which is why they got sued.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#656
post #114

Earlier quoted context omitted.

So, have you tried finding out what the programming was in October 1994? Or what cultural ephemera appeared in the TV guides of that era alongside the schedules? Either there's a copy for the week you want in an archive, or somebody's got one for sale, or most often neither. This can piss you off, if as it happened you had a reason to care.

To play devils advocate, completely on the terms of your argument, would it be better for that particular human artifact to be shredded and its contents melted into an anonymized data pool, or for it to exist in a museum archive, in its original form, such that future generations can better understand what it was like to be alive in 1994? I’d personally choose the latter, especially given that the 1994 tv guide is no…

A reasonable opinion, but I'd personally strongly choose the former.

A physical book in a museum archive is useless for 99% of the worlds population even if they really wanted that specific book and were able to find it, as they'd have to arrange for access, then travel (at incredible expense) to access it.

Maybe they could ask the museum to digitize it, but that's still going to be days of delay and tens of dollars of cost to access parts of that book, if the museum even offers that service. If we go slightly beyond your "melted into" statement, the chances of the book becoming useful to the public are much higher in the AI company's digital archive, which might turn into something like what Google Books could have been, given the right incentives and copyright law changes.

And of course that presumes that the TV guide is going to stay in the museum rather than been thrown out as part of curation (or realistically, long before it makes it into a museum). Neither museums nor archives hoard everything, throwing stuff out is - as far as I know - one of the key jobs of an archivist. And a 1994 TV guide, while useful to understand what it was like to be alive in 1994, likely doesn't contain much unique information. You don't need that specific guide.

If there are 52 weekly editions, of 10 different guides, you would likely get most of what you want from any one of them. And for the parts that you wouldn't - there's a good chance that you'll have a much easier time getting the essence of this knowledge from the anonymized data pool that all the content was melted into, rather than chasing 10 different museums to find the original magazines.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#657
post #579

Earlier quoted context omitted.

>Hopefully you can infer how this tracks to the written world and primary sources for language, technical manuals etc etc. I honestly can't and I think you can't either or you would have used an example with books/printed media rather than film, an entirely different ballgame.

OK... I'm going to assume good faith even though your wording makes it somewhat unlikely. Similar textual examples would be any text containing actual language as it's spoken in a given place or time. Or any factual textbook detailing the buildings present in a given location. Or any text book detailing a now defunct construction process. Or any text book (generally small run) detailing a niche interest, now missing…

[deleted]

Re: AI companies destroy physical books – let's scan rare books before it's too late

#658
What was worse: Putting all of our communication since ~2010 into a commercial walled garden? Or some books that were lying around in some bookstores or whatever (i.e. that nobody was interested in owning so far)?

And about what topic have I heard more complaints in the last 15 years (although the latter topic is just a few months old)?

Why is that?

If you say that I'm indeed wrong, and the latter one IS indeed much more important, then please tell me why? What is wrong with me then?

Re: AI companies destroy physical books – let's scan rare books before it's too late

#659
A few points for people: 1. some books are out of print. 2. some books CANNOT return to print. 3. all books prior to the 21st century are products of human minds. 4. copyright extends over the vast majority of printed material due to acceleration of literacy and printing access. 5. not all people value all books equally. 6. most books have a degree of historical interest (even cookbooks, which can say a lot about the economic health of a region when it is printes. culture is also clearly encoded in them). 7. many books are already lost, and historians are the ones who most voice the harms. 8. when a book is absorbed into the machine, it may remain vaugely accessible, but only on the good grace of the ones who pilfered it. 9. if no existant copies remain, then the price for access becomes effectively infinite. 10. removal of books denies human agency over access to information. 11. costs will follow a steepening curve much as ram did.

first they came for cookbooks, but i was no chef so i said nothing. second they came for handicraft, but i do not toil with fabrics or glue. next they came for homesteading, but i loathe the outdoors life. after, they came for biography, memoirs, and letters, but i am bored by the dead. finally, they came for my own little little interest, but nobody was left who appreciated books, so they too were ripped to shreds.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#660

I don't see any mention of Project Ocean - AKA Google books. Before AI they endevoured to digitize books in a massive online library. This inccluded rare and out of print books many which are archived at libraries. Because they had to preserve the books and return them in the condition they received then they created elaborate technology to accomplish this. The project was met with significant legal challenges from a…

> The project was met with significant legal challenges from authors and publishers which was eventually overcome

I don't think they were overcome. As far as I remember Google couldn't make the books available so they abandoned the project. They possess the scans (if they didn't delete them) but they won't be made public.

Post reply on HN