Live data from Hacker News

AI companies destroy physical books – let's scan rare books before it's too late

annas-archive.gl

811–820 of 961 posts

Re: AI companies destroy physical books – let's scan rare books before it's too late

#811
post #733

Earlier quoted context omitted.

Given all the bad PR around this issue - you'd think they'd do the 10 seconds of examination "is this rare" before putting it in the cutter. I'm genuinely surprised they don't.

The price is more or less an indicator. If they can shred a $1000 book, kudos to them or their budget. The idea they're simply shredding human history is too one-sided and alarmist. There is a huge long tail of rare-ish old and scrappy books out there that are definitely not the last copy of anything. That said, I'm pretty sure a very small % of these books fall into the mistake category. But life!

I've been noticing a lot of alarmist stuff in the anti-AI community, too.

I love how Oregon showed how to solve the electricity problem and is being entirely ignored. THEY'RE STEALING OUR ELECTRICITY!!!

Re: AI companies destroy physical books – let's scan rare books before it's too late

#812
post #734
post #639

Earlier quoted context omitted.

Knowing and enacting are two distinct things, "you should do what I want you to do" means you're already doing the wrong thing, and yes, there is an agreement on what the right thing is, its what is ethical, that is, what non profits and archivists are forced to do.

I can’t believe that in the multipolar world of 2026, where hundreds of conflicting world views coexist, where universalism is being disproven on a daily basis, it’s still possible to read things like “there is agreement on what the right thing is”. This is as blatantly false as claiming that the Earth is flat, and the fact that there is no such agreement (descriptive moral relativism) has been firmly established in…

If you think ethical arguments like "murder is bad" or "human knowledge should be preserved" and a demonstrable falsity like "the earth is flat" are equivalent arguments then you are so far gone you might as well be a flat earther.

Blind futurists and AI cheerleaders scare me on how cavalier and how many crimes against humanity they ignore.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#813

Earlier quoted context omitted.

Somehow I don't think they're looking for the books that have been copied over and over.

Why do you think that? Do you have evidence, or is this just a hunch? Why would Amazon waste money on uber rare books when there are thousands and thousands of not-so-rare books that could serve the exact same purpose?

For one they've already used the entire library genesis. Anything not in there is going to be obscure in some capacity.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#814

Earlier quoted context omitted.

That was due to negligence. Now it is profitable to keep the books private, and actively deny others access to them.

Ah, yes, because we all know OpenAI and Anthropic are famously just front operations for the high-end antiquarian book trade.

That's an in house operation that benefits them, not a front.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#815
post #801
post #773

Earlier quoted context omitted.

If it isn't scanned destructively, is it a liability? Can you do anything else with the book? What are its costs for storage in a way that retains the value of the book? If the assets of the warehouse are sold to another company (see also https://paizo.com/blog/paizo-restructuring-a-difficult-updat... ), what are your obligations for the format shifted copy that you retain? These questions imply that there's a liabil…

Questions don't imply anything, you're just asking things. I don't have the answers to your questions, but the point is that Google was able to scan books without destroying them while still avoiding legal consequences (with some effort).

The physical books that were scanned by Google were returned to the libraries.

https://btaa.org/library/programs-and-services/book-search/f...

> Will scanning harm the books?

> No. Google developed innovative technology to scan the content without harming the books. Any book deemed too fragile will not be scanned by Google, but may be treated by expert library staff. Once scanned, all print volumes are returned to the library collections.

That was an inherently different goal (borrow the books from the library, scan them, and return them) than the Bartz v. Antrophic ruling.

https://cases.justia.com/federal/district-courts/california/...

> Storage and searchability are not creative properties of the copyrighted work itself but physical properties of the frame around the work or informational properties about the work. See Texaco, 802 F. Supp. at 14 (physical), aff’d, 60 F.3d at 919; Google, 804 F.3d at 225 (informational); Sony Corp. of Am. v. Universal City Studios, Inc. (“Sony Betamax”), 464 U.S. 417, 447 (1984) (rightful interests). In Texaco, the court reasoned that if a purchased scientific journal article had been copied “onto microfilm to conserve space, this might [have been] a persuasive transformative use.” 802 F. Supp. at 14 (Judge Pierre Leval), aff’d, 60 F.3d at 919 (reducing “bulk[ ]” “might suffice to tilt the first fair use factor in favor of Texaco if these purposes were dominant“). In Google Books, the court reasoned that a print-to-digital change to expose information about the work was transformative. Google, 804 F.3d at 225 (Judge Pierre Leval). And, in Sony Betamax, the Supreme Court held that making a recording of a television show in order to instead watch it at a later time was copying but did not usurp any rightful interest of the copyright owner. 464 U.S. at 447, 455. Important to the Supreme Court’s reasoning was the expectation that most such copiers would not distribute the permanent copies of the work. Finally, in A&M Records, Inc. v. Napster, Inc., our court of appeals recognized the reasoning just explained, and therefore rejected by contrast a digitization effort that was touted as space-shifting but in fact resulted in the multiplication of copies shared with outsiders through a file-sharing service. 239 F.3d 1004, 1019 (9th Cir. 2001), aff’g in this part 114 F. Supp. 2d 896, 912–13, 915–16 (N.D. Cal. 2000) (Judge Marilyn Hall Patel) (citing Sony Betamax and Texaco).

> Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company. This use was even more clearly transformative than those in Texaco, Google, and Sony Betamax (where the number of copies went up by at least one), and, of course, more transformative than those uses rejected in Napster (where the number went up by “millions” of copies shared for free with others).

---

The AI training isn't borrowing from libraries and returning from libraries. Instead, it is buying a book, format shifting, and retaining that format shifted version from its own use. The company can't do anything else with the book once they've format shifted it. They can't donate it and they can't resell it. In that light, destructively scanning the book is the best option. There is no value in trying to non-destructively scan it because otherwise all it would do is sit in a warehouse and cost money to pay for storage of something they can't sell.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#816

Earlier quoted context omitted.

Isn't this what Dario does?

If Dario appeals to feeling, yes. Provide more context if you want a better feedback not whataboutism.

Dario keeps rambling about how China is bad, like the US doesn't do those things itself. That appeals to American exceptionalism which every American is taught in school, it's just not taught explicitly because it's ingrained in the very lifestyle that Americans have

Re: AI companies destroy physical books – let's scan rare books before it's too late

#817
post #807

Earlier quoted context omitted.

Copyright protection is automatic. Publishing of copyrighted material requires that it be deposited with the Library of Congress.

Right, that's why I said we should make copyright require deposit and registration (like it used to). Publish your work without its copyright ID for people to use to reference the LoC database? It is now public domain.

That would require a renegotiation of the TRIPS Agreement and the Berne Convention with the rest of the countries of the WTO.

https://www.wto.org/english/tratop_e/trips_e/ta_docs_e/modul...

    (iii) Automatic protection A key feature of the Berne Convention, and thus also of the TRIPS Agreement, is that copyright protection - unlike most other forms of IPRs - may not be subject to any formality of registration, deposit, or the like. This principle is contained in Article 5(2) of the Berne Convention, that has been incorporated into the TRIPS Agreement.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#818
post #729
post #508

Earlier quoted context omitted.

Like all things, Google will eventually realize they cannot make significant ad revenue and they will eventually give up and discontinue serving this, though I doubt it's more than a scratch in terms of disk space. It's great they did this, but the Google that is today cannot be trusted with data of public value anymore.

I think you're missing the point. Maybe I'm wrong, but I'm pretty sure they're trying to point out that this book scanning doesn't need to be destructive. I'm not sure if AI companies are using a scanning method that damages the book or not, but they destroy the books after scanning to avoid copyright issues (ie they aren't duplicating the books). This Google project seems to demonstrate that this isn't actually nece…

If you're going to destroy the book anyway you'll go with the easier method for a good scan(separate the pages from the spine)

Re: AI companies destroy physical books – let's scan rare books before it's too late

#819

Earlier quoted context omitted.

> permanently locking human knowledge inside private corporate servers History tells us that very few "permanent" situations are truly permanent. Provided a set of information has value (which in this case it clearly does) then the overwhelming likelihood is that, eventually, through some method or other, the information will become public.

> History tells us that very few "permanent" situations are truly permanent. If you destroy the only copy of a physical artifact, the situation is as permanent as it can get.

I mean books get destroyed all the time, really they are a major pain in the ass to keep together, especially as they age. Paper loves to crumble. Insects think they are tasty. Floods and fire love destroying them too.

So physical books are rather non-permanent themselves.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#820
post #810

Earlier quoted context omitted.

This is already required for new books published in the US, and has been for more than one hundred years. It’s called “mandatory deposit”

Case law seems to be that mandatory deposit is unconstitutional, fwiw. https://en.wikipedia.org/wiki/Valancourt_Books_v._Garland

That revolves around print on demand for books that are out of copyright or where the copyright has been abandoned.

> Background Valancourt Books is a print-on-demand independent publishing house specializing in rare and out-of-print books. Valancourt had not registered its books for copyright as the Library of Congress already had original-edition copies of the books Valancourt republishes and any new material in its publications was limited to notes and introductions.

If you want a physical copy of The Sorrows of Satan, you can buy it from them.

Their argument is that the Library of Congress already has a copy of the book ( https://search.catalog.loc.gov/instances/a0f8fcfe-a255-55d2-... ) and having them deposit it again would be unnecessary.

> The Copyright Office has stated that it would modify the language of its deposit demand letters and withdraw its demand for copies if the Copyright Office was notified of the copyright's abandonment.

> Several legislative changes have been proposed to address all elements of the case: changes to Section 407 to tie some legal benefit to the deposit, monetary compensation to copyright holders for depositing books, and regulation for a simple and costless method of copyright abandonment.

That doesn't change that if you were to publish a book today (or for that matter, have published a book in the past 100 years in the US), you are required to deposit a copy of the book with the Library of Congress.

Post reply on HN