Live data from Hacker News

AI companies destroy physical books – let's scan rare books before it's too late

annas-archive.gl

531–540 of 961 posts

Re: AI companies destroy physical books – let's scan rare books before it's too late

#531
post #508

Earlier quoted context omitted.

Like all things, Google will eventually realize they cannot make significant ad revenue and they will eventually give up and discontinue serving this, though I doubt it's more than a scratch in terms of disk space. It's great they did this, but the Google that is today cannot be trusted with data of public value anymore.

They are also an AI company now. Why would they stop?

They could use them as training data, without providing access the actual books.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#532
post #508

I don't see any mention of Project Ocean - AKA Google books. Before AI they endevoured to digitize books in a massive online library. This inccluded rare and out of print books many which are archived at libraries. Because they had to preserve the books and return them in the condition they received then they created elaborate technology to accomplish this. The project was met with significant legal challenges from a…

Like all things, Google will eventually realize they cannot make significant ad revenue and they will eventually give up and discontinue serving this, though I doubt it's more than a scratch in terms of disk space. It's great they did this, but the Google that is today cannot be trusted with data of public value anymore.

It's data for their AI pipeline. Basically digital gold.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#533
post #508

Earlier quoted context omitted.

Like all things, Google will eventually realize they cannot make significant ad revenue and they will eventually give up and discontinue serving this, though I doubt it's more than a scratch in terms of disk space. It's great they did this, but the Google that is today cannot be trusted with data of public value anymore.

They are also an AI company now. Why would they stop?

Also, why would they ever share their collection?

Book scans, secreted away, are worthless to the public.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#534
post #464

Earlier quoted context omitted.

It was never that dramatic, and it’s declining day by day. It’s a panic over a real but small problem. Order of magnitude more water is lost from wasted irrigation (e.g. during rain, of fallow fields, sprayed into windy air, etc) than data centers.

Problem with data centers is that companies want to build them near densely populated areas that already have problems with water supply and high utility bills.

1. They don't HAVE to use water. Air cooling, closed loop cooling, waste-water cooling, and so on, are options. Easy to regulate. Evaporative cooling is more energy efficient though, but a complete non-issue in places with abundant water and a non-option elsewhere.

2. Datacenters have been shown to reduce utility prices. They provide suppliers with previsible long term demand which allows for cost-effective network and production planning.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#535
post #299

I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collective responsibility. Thus, the preservation of this knowledge, its availability, and accessibility are t…

> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.

How can anyone say this with a straight face. The knowledge is not destroyed, it is transformed. You can make use of it today in the form of LLMs and the scans still exist. Nothing was lost. It's literally no different from them buying books and stocking them in a private library not open to the public. It's not called the Scanning of Alexandria because if it was, it wouldn't have made a blip in the history, Alexandria's libraries were burned, those books, that knowledge was destroyed. Then only thing being destroyed here is physical copy (again for the people in the back: a copy).

> Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available.

Those same copyright restrictions are exactly what would prevent them from sharing the archives. Your beef is with copyright, not the AI companies who are (in this one, rare, instance) following copyright laws/rules.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#536
I'm a little confused about the value of scanning rare books since this story came out. I have a small collection of "rare" books and they're not really bounties of information, at least not modern information. I know novel training corpus is important but the information in rare non-fiction books is commonplace or out-dated. And the information in old rare fiction-books are originals for which reprints exist or just uninteresting stories that aren't really worth anyone's time except collectors'.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#537

The main question is why aren't they leaking it to AA themselves? Trying to keep an edge with their training sets? Isn't it ridiculous, considering the sheer size of them and statistical insignificance of the set differences?

> The main question is why aren't they leaking it to AA themselves? How is this a question at all? They’re scanning books because the courts determined that it’s the only way to use that data. They are forbidden from using digital copies found on places like Anna’s Archive. They must acquire and scan the book. They cannot redistribute the book. The Internet Archive tried that and the courts shut it down. You cannot s…

> You cannot scan a book and share it without violating copyright law.

Hence the “leaking” part.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#538

Earlier quoted context omitted.

You can't compare an island or food to books though, because the former are only really valuable in their atomic form. Like if you take detailed pictures, people can't live on or eat those pictures. You scan books though and all the value they provide is still fully available in bit form, because books merely contain information. And of course being in bit form means they can be trivially copied, technically. But the…

It's culture. These scans are of course not distributed, so the value doesn't quite continue to exist. You say "culture", but is is culture , and culture that was interesting enough to write down. Writing a book in the 70s isn't a trivial thing, there are gatekeepers, editors, publishers etc., even for your little book about local history, and the information is valuable to the community it's about. It affects their…

The physical work before was never distributed either (which is what makes them economically valuable), so the the point is mostly moot. I say mostly because at least when these neglected works become part of a training corpus, the knowledge they contain can be surfaced on demand, or even by accident. Think of it like the grandpa telling stories of his experiences to the grandkids, to the best of his recollection (and imagination), going on related tangents as they occur to him or are triggered by the grandkids' questions, etc.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#539

Earlier quoted context omitted.

>From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge. How tho? Ever sunday at the flea market, I see thousands of books that are rotting, hoping for someone to buy them or at least take t…

Yep, this feels pretty much it. Looking at the "rare books", it was books that nobody would care about or would just rot away anyway. (The example book of Old books of agriculture is probably not that important today) Liberians have to accept that much of their job is sending books to be burned. A lot of them try their best to get people to be interested in older books, but they have to make way for "newer" books ins…

Well if you're appealing to librarians:

"Hey, guys, when the librarians get pissed about the destruction of books, it’s time to put those listening ears on.

Because we are very comfortable with the idea that books are tools that can be retired. What’s happening right now is not that...."

https://bsky.app/profile/annabookwriter.bsky.social/post/3mt...

Re: AI companies destroy physical books – let's scan rare books before it's too late

#540
post #461

Earlier quoted context omitted.

So if I scan a book, sell it, and keep using the scan, is that legal? (Spoiler: That's not legal. It's a violation of IP law.)

Selling it is not allowed. The inability to sell it does not require its destruction.

Why would a company keep the hard-copy around at the risk of it being inadvertently given away, resold, etc.? It's a huge outstanding liability given that the illegal copying of works -- the other part of that case -- is what they settled out of court for some huge amount of money. Destruction is the only thing that makes sense.

I'm old enough to have been around when DCMA legislation was under discussion. Many people were dead-set against it and raised concerns over matters exactly like this. In Rainbows End (2006), Vernor Vinge wrote about a similar scenario where a robot went through the university library shredding books, and scanned the shredded pieces to recombined them into a digital archive.

Anthropic may be doing shady things and may have even done this on their own recognizance, we just don't know. As it stand, this is 100% a consequence of US copyright law, much of which was written by large corporations to protect their own assets.

Post reply on HN