Live data from Hacker News

AI companies destroy physical books – let's scan rare books before it's too late

annas-archive.gl

661–670 of 961 posts

Re: AI companies destroy physical books – let's scan rare books before it's too late

#661
post #609

Earlier quoted context omitted.

Why would a company keep the hard-copy around at the risk of it being inadvertently given away, resold, etc.? It's a huge outstanding liability given that the illegal copying of works -- the other part of that case -- is what they settled out of court for some huge amount of money. Destruction is the only thing that makes sense. I'm old enough to have been around when DCMA legislation was under discussion. Many peopl…

I agree fully that it makes logistical sense. But it is not a legal requirement, and they should not be permitted to use that as an excuse to wash their hands of their own decisions.

I think I agree with you in spirit. I don’t like what these companies are doing, and Anthropic’s actions can for the most part stand on their own. Copyright law is just a special interest of mine, and I do think it’s important to recognize what external incentives exist and what they prioritize. Because other companies will act in similar manners under the same incentive structure, and the problem is going to cascade and magnify if it hasn’t already. There are active court rulings setting precedence for this behavior - take note!

Re: AI companies destroy physical books – let's scan rare books before it's too late

#662

Earlier quoted context omitted.

A small change in the copyright law would fix this problem. Something like: If a company is scanning material protected by copyright, it has to send a digital copy of the scanned material to Library of Congress within 5 working days.

The Library of Congress already has a copy of every book published in the US. How would this help?

If the Library of Congress has a digital copy, it would be easier for them to distribute the work after the copyright of the work expires. That would be a public benefit.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#663

Earlier quoted context omitted.

Is there a difference with regard to the metaphor?

Depends on the point being made about “historical precedent” and the lessons to be drawn from such. Also, helps to clarify what exactly the commenter was referring to and possibly help distinguish the centuries-spanning decline of the Library of Alexandria from the violent fate of the Serapeum.

My take was generally: the loss of Alexandria's collection represents a calamitous loss to our collective culture.

(But my knowledge of Alexandria extends only to episodes of "COSMOS" and "Connections").

Re: AI companies destroy physical books – let's scan rare books before it's too late

#664

Earlier quoted context omitted.

>Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available. Aren't the copyright laws forcing them to do this the very ones that would make such archives illegal? The books that could be in such an archive are the books that don't need to be destroyed.

Copyright law eating itself…

no matter how you look at it, this is a systemic failure. if as a society we're going to mass scan our history then we should be building an archive for the future. not using availability of information as a moat. not doing it over and over again and throwing it away because of some odd rules to protect someones market position. not using it as an excuse to put paywalls around 80 year old field guides to field rodents in western massachusetts. not taking texts that had limited value and mining them for turns of phrase to be piled up into a useless grey goo.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#665
I keep seeing headlines, videos, etc and the recent copyright court case, Anthropic v. Bartz (1.5 billion dollars) gives the best context around this. I encourage everyone to read the full thing, but here are some excerpts:

> Anthropic spent many millions of dollars to purchase millions of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals. Each print book resulted in a PDF copy containing images of the scanned pages with machine-readable text (including front and back cover scans for softcover books). Anthropic created its own catalog of bibliographic metadata for the books it was acquiring. It acquired copies of millions of books, including of all works at issue for all Authors. Anthropic may have copied portions of Authors’ books on other occasions, too — such as while copying book reviews, academic papers, internet blogposts, or the like for its central library. And, Anthropic’s scanning service providers may have copied Authors’ print books along the way to delivering the final digital copies to Anthropic. But neither side here specifically raises legal issues implicated by any such copies. Nor will this order

Also the summary:

> To summarize the analysis that now follows, the use of the books at issue to train Claude and its precursors was exceedingly transformative and was a fair use under Section 107 of the Copyright Act. And, the digitization of the books purchased in print form by Anthropic was also a fair use but not for the same reason as applies to the training copies. Instead, it was a fair use because all Anthropic did was replace the print copies it had purchased for its central library with more convenient space-saving and searchable digital copies for its central library — without adding new copies, creating new works, or redistributing existing copies.However, Anthropic had no entitlement to use pirated copies for its central library. Creating a permanent, general-purpose library was not itself a fair use excusing Anthropic’s piracy.

https://copyrightalliance.org/wp-content/uploads/2025/06/Bar...

Re: AI companies destroy physical books – let's scan rare books before it's too late

#666

Earlier quoted context omitted.

I'm only one person, but I scan old books that had an impact on me growing up, and upload them to archive.org. Thankfully there are others that do the same. (And to be sure, FWIW, these are books that have not been printed for about 50 years—I suppose the software community would call them abandonware.)

If they're 50 years old they're young, and archive.org will likely block access. If they're not already on annas-archive (or the copy there is trash), your best bet is an anon upload to libgen.

Thanks.

There was a time of course when you could pull my books down from archive.org as PDFs. Perhaps that time will come again.

I'll look into libgen.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#668
post #579

Earlier quoted context omitted.

>Hopefully you can infer how this tracks to the written world and primary sources for language, technical manuals etc etc. I honestly can't and I think you can't either or you would have used an example with books/printed media rather than film, an entirely different ballgame.

OK... I'm going to assume good faith even though your wording makes it somewhat unlikely. Similar textual examples would be any text containing actual language as it's spoken in a given place or time. Or any factual textbook detailing the buildings present in a given location. Or any text book detailing a now defunct construction process. Or any text book (generally small run) detailing a niche interest, now missing…

[deleted]

Re: AI companies destroy physical books – let's scan rare books before it's too late

#669
post #352

Earlier quoted context omitted.

Nowhere in this description did it require destruction of the physical book. This is being done because it's easier to scan a shucked book, and this explanation is circulating because it's easier to blame it on the law and that pesky meddling government.

So if I scan a book, sell it, and keep using the scan, is that legal? (Spoiler: That's not legal. It's a violation of IP law.)

Probably, if you sell it after the copyright expires.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#670
post #299

I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collective responsibility. Thus, the preservation of this knowledge, its availability, and accessibility are t…

A small change in the copyright law would fix this problem. Something like: If a company is scanning material protected by copyright, it has to send a digital copy of the scanned material to Library of Congress within 5 working days.

IANAL, but I recall copyright law being far too complex for any easy Protected/Not Protected test to exist.
Post reply on HN