Live data from Hacker News

AI companies destroy physical books – let's scan rare books before it's too late

annas-archive.gl

771–780 of 961 posts

Re: AI companies destroy physical books – let's scan rare books before it's too late

#771

The scale of problem seems a bit overblown. Anna's Archive paint a picture like AI companies are some movie villains burning books so no one can see them. but in reality they just disassemble them into pages because it's cheaper and faster to scan. Most of these books is highly specialized, they been collecting dust on shelves for decades and nobody need them. But the problem is real. Even if these books aren't neede…

> You can't to ask a neural net to give you a exact copy of a page from that book

That's the problem. I tried the other day to find the verbatim quote from one of the Gabriel García Márquez's books - a single effing sentence and I nearly lost my mind - every LLM would be like "this is a copyrighted material and I can't share it verbatim". WTF? These books are widely known, published in many languages, and yet no amount of money spent on tokens can give me the exact fucking sentence, just like the author intended? Really? I need to find, buy, download and search through the actual book to quote something that Márquez wrote at some point? What is even worse is when the LLM just outright lies to you, giving you the quote, but slightly rephrased.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#772
post #130

“Rare books” usually refers to rare editions of books. Any books out there where there are only a few extent copies of the text itself, are probably not of very much interest or social value, since almost no one is able to read them, by definition. If you think there is priceless knowledge locked up in books so rare that it is on the verge of being lost forever, then AI labs are not really the problem!

The object itself is the valuable thing, not necessarily the text. Why are so many in the discussion overlooking something so obvious?

Re: AI companies destroy physical books – let's scan rare books before it's too late

#773
post #729
post #508

Earlier quoted context omitted.

Like all things, Google will eventually realize they cannot make significant ad revenue and they will eventually give up and discontinue serving this, though I doubt it's more than a scratch in terms of disk space. It's great they did this, but the Google that is today cannot be trusted with data of public value anymore.

I think you're missing the point. Maybe I'm wrong, but I'm pretty sure they're trying to point out that this book scanning doesn't need to be destructive. I'm not sure if AI companies are using a scanning method that damages the book or not, but they destroy the books after scanning to avoid copyright issues (ie they aren't duplicating the books). This Google project seems to demonstrate that this isn't actually nece…

If it isn't scanned destructively, is it a liability?

Can you do anything else with the book? What are its costs for storage in a way that retains the value of the book? If the assets of the warehouse are sold to another company (see also https://paizo.com/blog/paizo-restructuring-a-difficult-updat... ), what are your obligations for the format shifted copy that you retain?

These questions imply that there's a liability that exists when retaining the original that has little value to the company. And they (the books) aren't assets that can be resold.

It's easier (and cheaper), has no ongoing costs for physical storage, and answers those questions without creating legal entanglements for the future company.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#774
post #299

I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collective responsibility. Thus, the preservation of this knowledge, its availability, and accessibility are t…

The idea that these corporations or any of the literal sociopaths that work for them give the slightest bit of a shit about "benefitting humanity" is hilariously naive. The one and only thing these entities care about is money, and making as much of it as they can. If they could get away with it, they'd commit every crime that exists if it meant they get a quarter of a percentage increase in their quarterly earning reports.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#775
post #733

Nondestructive scanning can cost 10x as much. This is about cost. It is not about preservation. Google never destroyed the books it scanned. Amazon and Anthropic are attempting to save money. They are not considering whether or not a book is rare. They are treating books as a commodity. Rare books are rare. It is easy enough to identify when there are a limited number of copies of a book. The issue is saving money on…

Given all the bad PR around this issue - you'd think they'd do the 10 seconds of examination "is this rare" before putting it in the cutter. I'm genuinely surprised they don't.

The price is more or less an indicator. If they can shred a $1000 book, kudos to them or their budget.

The idea they're simply shredding human history is too one-sided and alarmist. There is a huge long tail of rare-ish old and scrappy books out there that are definitely not the last copy of anything.

That said, I'm pretty sure a very small % of these books fall into the mistake category. But life!

Re: AI companies destroy physical books – let's scan rare books before it's too late

#776

Earlier quoted context omitted.

Depends upon what you want. For example, knowledge about how the book smells when you open it, something that reader do talk about enough I don't think this should be a strawman, is lost. But, that is about the experience of reading the book, not the knowledge of the book. The exact text? Yeah, I think that is largely lost as well. This is a summary. And for rarer books, it will be a particularly bad summary. The bas…

> But, that little bit of data is a bit more data than existed before, No, it’s not. Undigitized data is still data. Wording, style, nuance, type, artwork, binding, metadata, attributions, citability… all of these things are permanently lost, which is a fucking tragedy because they don’t have to be. Even if you’re not willing to take the time to scan every page in a cradle scanner, as rare books should be, (and don’t…

>all of these things are permanently lost

A small fraction of them is saved in the model. Far more is saved in the digitized copy as long as they keep it which they have plenty of incentives to do so (future training of newer models).

That's more than what happens if that book was burned or sent to a landfill, but less than if the book is giving a loving home.

>They removed the pages from the binding, scanned them on a high speed conveyor belt scanner which yielded full color 600 DPI jp2 images, placed the pages back in the binding like a folio that could be re-bound if needed, vacuum sealed them, and stored them in a salt mine.

My understanding is that this simply isn't legally allowed for these books. The original must be destroyed for the digital copy to not be copyright infringement.

>That’s a false dichotomy.

I pointed out there is a spread of possible outcomes and that different people are considering different outcomes and the comparison of if this is good or bad depends upon which outcome one considers. I even mention that both outcomes are sometimes right. That's about as far from a false dichotomy as I can see it.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#777
post #745

Earlier quoted context omitted.

> Google providing what would have been one of the greatest storehouses of readily available knowledge in the world I'd be worried about how much they'd be charging for access once they had the monopoly on so many rare books.

Then at least there would be outrage to drive the passing of the needed legislation which otherwise hasn't come to pass anyway.

Copyright law desperately needs a production requirement or allowance.

The copyright owner must make new copies of the work available; the price must be no greater than the original price (not inflation adjusted). And if they fail to do so, anyone may produce copies and escrow the original price (less the cost of production) for collection by the copyright holder.

That means that orphan works are effectively in the public domain. Calculus professors can ask students to get the cheaper 2nd edition, not the latest 22nd edition. And a company like Google could make scanned works available in their entirety for a small amount of money for each work. And the copyright holder still gets their end, without having to arrange a printing or hold stock.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#778

Earlier quoted context omitted.

Puts the burden on government to store what is probably 90% worthless material. Copyright should really be amended so that once out of print and a grace period it’s free use. I am probably more of an anarchist in this regard. Similar to my belief that anyone should be able to ingest any data you put online, once a book is no longer being print it should be able to be used for commercial or personal use for free. Simi…

> Puts the burden on government to store what is probably 90% worthless material. That's what governments are for.

That's what governments already do.

https://www.copyright.gov/mandatory/

> All works under copyright protection that are published in the United States are subject to the mandatory deposit provision of the Copyright Act (section 407 of Title 17).

> This law requires two copies of each work published in the United States be deposited with the Copyright Office within three months of publication. Works deposited under this law are for the use of the Library of Congress. Usually, deposited copies must be the “best edition” of the work, which means they must conform to the Library of Congress’s preferred specifications.

> Mandatory deposit applies to any work published in the United States. This requirement does not apply to works first published in a foreign country until they are published in the United States. Copyright registration is optional, but it provides additional legal benefits and fulfills the mandatory deposit requirement with the submission of the required copies.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#779
post #709

Earlier quoted context omitted.

It's possible for a datacenter to use scarce well water irresponsibly. They don't have to, and the vast majority don't. The lie and moral panic is that all datacenters necessarily waste precious drinking water; it's patently false and used by agitators to push, unwittingly or not, a Chinese Communist Party agenda.

> a Chinese Communist Party agenda is it actually? it feels more like US propaganda that we have to let our oligarchs run roughshod over us because of what we imagine the big bad CCP might want. i dont think the CCP cares whether there's data centers in rural america. Regulation that requires closed loop cooling seems simple enough, same with lots of the other problems people have with data centers: * sound and infra…

> i dont think the CCP cares whether there's data centers in rural america.

They certainly care that the US loses the race for AI.

Post reply on HN