Live data from Hacker News

AI companies are shredding rare books

twitter.com

191–200 of 559 posts

Re: AI companies are shredding rare books

#191

I've limited sympathy for the publishers. It pisses me off to reflect that they can sit on works until copyright expires, keeping them out of print. There's no real need for any of these so-called rare books to be rare while they're under copyright. And related to this, the books that are in print are mostly only in print in the shittiest way. I often see well-made books from the 17th or 18th centuries which are stil…

Books in the 17th and 18th century often didnt get bound by the publisher. They were done by independent binders for _custom_ orders. You’d see whole libraries with the owners binding/cover standards rather than per book. Books in that time were _luxury_ goods. Most people could not afford them. One of the ways that was changed was to introduce cheap, mass produced bindings that were lower quality than the bespoke ar…

One of my relatives is one of those micro publishers. Selected works are printed and bound to extremely high standards and materials in a way that their customers are willing to pay 1k-25k+ per book. These editions only have a couple of prints and are mostly made to customer requests.

There is a market for these type of books, albeit a very small one.

Re: AI companies are shredding rare books

#193

Earlier quoted context omitted.

It’s been determined that training on lawfully acquired works is fair use. Presumably in this discussion of shredding physical books Dario and Sam are not pulling heists at the local library. I’m sure there’s ongoing litigation, and better sources than this, but fair use was determined in June 2025 in a sf federal district court https://www.goodwinlaw.com/en/insights/publications/2025/06/... Similar conclusion vs met…

You don’t know what “Fair Use” is. Fair Use is not an activity that you engage in. Fair Use is not a category with criteria that you meet. Fair Use is not a precedent that paves the way for everything afterwards. Fair Use is a defense that can be used in court when you’re named in a copyright lawsuit. Fair Use is how you justify your actions before the court finds infringement.

>>It’s been determined that training on lawfully acquired works is fair use

>You don’t know what “Fair Use” is.

This isn't the opinion of some armchair HN commenter. Actual judges have affirmed this, as other commenters in this thread has pointed out.

Re: AI companies are shredding rare books

#194

you wouldn't believe how much shredding your local library does in the name of space conservation and to address changing borrower preferences and demographics. I doubt any AI company comes close to the annual combined library turnover

Where I live there's a large library-associated biannual book sale where books are (effectively) reverse auctioned over a period of a few weeks. At the end, anything (with a few exceptions, like the Collector's Corner) that isn't sold is disposed of, with a large 18-wheel truck sized dumpster filled with items to be sent for pulping and recycling. The price at the end is $1 for a grocery bag full of books, so things that don't sell truly are perceived as worthless.

https://booksale.org/

Books are information delivery vehicles. We mostly shouldn't care about them any more than we care about a particular set of bits on a disk.

Publishers also pulp large numbers of books themselves. This is a consequence of the Supreme Court's Thor Power Tools ruling, which clarified tax rules in the US so that keeping large inventories of unsold books was less economical.

Re: AI companies are shredding rare books

#195

Maybe we can kill two birds with one stone: digitize rare books and reverse the damage from Authors Guild v. Google [1]. Let AI companies do this. But require them to make the digital copies public. Maybe with a multi-year delay, to give the original scanner advantage to doing it. [1] https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,... .

Yeah, I was thinking along these lines... Let's say one of the books to be digitized and destroyed is the sole remaining copy of a book from 1850, which is now considered public domain. On one hand, hoarding such a book, stealing its content from the public domain, locking its content behind a for-profit machine, and destroying the only remaining copy is clearly wrong. It's equivalent to stealing a public resource, j…

I think this misconstrus what public domain is.

It provides a freedom to circulate, but not access to the material. It is not a public owned resource.

Turning a copy over to the public or state might be an interesting requirement for obtaining a copyright, but instituting that fix for new works now would have a 70 year lag time.

Think of it this way, if I copyright a book and put it in my dresser for 70 years, that doesn't give the public the right to access it or come into my house and scan it after expiry

Re: AI companies are shredding rare books

#196
post #50

From what I understand, the rare books in question are not some historically relevant medieval manuscripts, but rather some relatively recent books (still under copyright) for which there are few print copies available for purchase. I am not sure that physically destructing one copy of this type of book to preserve its contents digitally is so bad. Pretty much anything that is still under copyright should be valuable…

Is slightly modifying some weights in a markov chain somewhere really preserving it's contents?

Re: AI companies are shredding rare books

#197
post #105
post #85

Earlier quoted context omitted.

This is really dishonest framing, unless you really, honestly can't tell a difference between pulping a mass market paperback romance novel that there's 3 million of in circulation, and shredding an 18th century botanical text that there's only 2 copies of in existence.

You call of dishonest framing, but you're begging the question twice. Once that AI companies are really shredding 200 year old rare books, and once that libraries are only weeding mass market pulp fiction.

That is literally what this article is about.

Re: AI companies are shredding rare books

#198

Earlier quoted context omitted.

Yes, actually you can. The fetish of physical book worship is a fossil of an age when information storage and retrieval was much harder.

I strongly disagree with this take, a book might remain readable thousands of years from now but very few if not zero of our digital data formats likely will. We shouldn't be so quick to throw away diversity in the way information which may be useful for future generations is stored for the long term.

Paper doesn't easily survive for thousands of years.

You know that pleasant used bookstore smell? It's paper slowly decomposing.

Re: AI companies are shredding rare books

#199
post #8

> You can reprint a bestseller. You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal. So it's going to accelerate. Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright? IDK about the shredding, it's not nice, but it's more a problem with copyright…

> You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal. An 18th century book would be out of copyright so why would it be illegal to keep the original and scan it?

[deleted]

Re: AI companies are shredding rare books

#200
post #162
post #97

Earlier quoted context omitted.

Not in a traditional sense, but obviously on the surface, they're using the info to regurgitate in some fashion and serve back. The point is that even under the best intentions, hallucinations occur. Then there's the fact that most models have an ideological bias programmed into them. The only expectation I have is for companies or anybody else to not destroy rare books. Is that such a tall order?

> they're using the info to regurgitate in some fashion and serve back. Sure, in the same sense that they regurgitate any other text they consume. LLMs by definition do not have the full training dataset available, though. It’s far larger than the resulting model. So they can’t reliably reproduce full text without an external source (or if it’s in the training data repeatedly). ChatGPT actually refused to give me a b…

> LLMs by definition do not have the full training dataset available, though.

That makes it even worse, then. This proves the original point.

> The idea that people or corporations should hold onto books forever because of a cultural “ick” about throwing out books is a bit ridiculous.

If we're building black and white straw man arguments, then sure, let's not archive anything.

Post reply on HN