Live data from Hacker News

AI companies destroy physical books – let's scan rare books before it's too late

annas-archive.gl

481–490 of 961 posts

Re: AI companies destroy physical books – let's scan rare books before it's too late

#481
post #299

I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collective responsibility. Thus, the preservation of this knowledge, its availability, and accessibility are t…

A small change in the copyright law would fix this problem. Something like: If a company is scanning material protected by copyright, it has to send a digital copy of the scanned material to Library of Congress within 5 working days.

Wrong end of the pipeline; we should instead demand digital copies of media be sent to the Library of Congress in order to obtain copyright, along with a registration fee to pay for indefinite storage and other costs. Registration should be mandatory if you want copyright. For things like books where a machine readable text format existed, it should be mandatory to include (so no requiring OCR). Access to the archive should be available for research use (including ML training) at cost.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#482

You ask "Why destroy physical books?" I ask "Why save physical books?" If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere. Owning and storing physical books is not free, there is a real cost. Even owning and storing the scans is not free, especially when IP rights and challenges get involved, since a scanned book with no distribut…

Massively missing the point. Having paper books isn't the point. Preserving copies of books is the point, so history isn't lost. Often those books only exist as paper copies due to their age. I don't really care what happens to the paper copies, only that their content is preserved in a way that is accessible. Anthropic's private servers aren't it.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#483

Earlier quoted context omitted.

A small change in the copyright law would fix this problem. Something like: If a company is scanning material protected by copyright, it has to send a digital copy of the scanned material to Library of Congress within 5 working days.

Wait, so the library of congress is suddenly responsible for probably petabytes a day of incoming scans? To what end? Do they have to index it and make it available? Do they have to check the accuracy and integrity of the scans? How does this help anything, except create more work to throw in the trash?

This is already required for new books published in the US, and has been for more than one hundred years.

It’s called “mandatory deposit”

Re: AI companies destroy physical books – let's scan rare books before it's too late

#484

What often gets missed is that they are buy one physical copy and turning it into a digital copy. They have done zero to destroy the durability. In fact, it’s probably more durable. If the physical copies are scarce, that is due to the publisher and copyright laws and not them buying and converting a single copy.

>that is due to the publisher and copyright laws

Agree. While it certainly isn't the most environmentally friendly to render huge stacks of paper into waste, the real issue is copyright creating scarcity (inability to copy the thing).

Re: AI companies destroy physical books – let's scan rare books before it's too late

#486
post #395

It is not a big deal. Since the invention of the printing press any important book has been duplicated by thousands, tens of thousands or even million of units. Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is stored somewhere) is not problematic for Humanity. By the way, I always search for second hand books. Most of the books there ar…

First of all, it IS destroyed and it is a big deal. Hardcover copies of books especially 1st - 2nd edition ones (even with mistakes) are rarer than digital scans. Maybe the Bodleian Library at Oxford University should give all their rare books to AI companies to scan and destroy them since it is not a "big deal" anyway. Except that when they did do a pilot with OpenAI to scan these rare books, [0] they did NOT destro…

Why wonder? The answer is abundantly clear if you follow the news. If a book is copyrighted under U.S. law, scanning and destroying counts as a format conversion which qualifies it as fair use, so there is no need to negotiate with copyright holders. See Judge William Alsup’s decision. If Anthropic did not destroy the books after scanning it would have not won the lawsuit, and scanning would be illegal. If a book is already out of copyright then of course they do not have to destroy it afterwards.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#487
I genuinely have yet to connect on this idea that we are “burning Alexandria” or AI companies are ruining the future of humanity because most of these books are absolutely junk.

I do think book copyright law needs a ton of work but I think most folks are simply taking their bias against AI and creating hyperbolic scenarios. I am sure there are some gems in the lot and I am equally certain they may be scanning dupes of the same material but even at scale I have a hard time seeing the significance. Most of these published work in the last 60 years is absolutely junk garbage. The good stuff usually has a lot longer run so more volume in circulation. You can go pick up lots of books that are 100+ years old for a couple bucks or cheaper because this stuff has no value.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#488

During World War II and its immediate aftermath, between 35 million and 40 million books were destroyed in Germany due to Allied actions

They did rather bring that on themselves though

But think of the books!

Re: AI companies destroy physical books – let's scan rare books before it's too late

#489
post #404

Earlier quoted context omitted.

Good to see someone making this point. I'm confused by the panic, because they are making it out like AI companies are destroying every copy of the book. They only need one, and they destroy it after scanning it only because they don't want to store them all. And storing or archiving all these books is not a trivial task.

They are destroying them because it’s easier to scan them if you slice the binding.

And very hard to re-assemble after you have done that. They are not in the book binding business.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#490

Earlier quoted context omitted.

They're not destroying rare manuscripts or incunables. They're destroying one (1) copy of a mass-produced item for each AI company. Public libraries destroy millions more yearly as a matter of routine. This is just part of a CCP-aligned moral panic, along with the water use nonsense, and similar with the soviet-aligned moral panic that destroyed the civil nuclear industry 40 years ago.

> This is just part of a CCP-aligned moral panic, along with the water use nonsense What was nonsense about water use?

That AI data centers are drinking up local ground water for cooling. It isn't (or wasn't) nonsense, though. It was/is a real thing, though it seems to be on the out in favor of closed loop cooling after the massive and still on-going public outcry.
Post reply on HN