Live data from Hacker News

AI companies are shredding rare books

twitter.com

371–380 of 559 posts

Re: AI companies are shredding rare books

#371

Earlier quoted context omitted.

Easier to send them through a duplex scanner (or put them in a flatbed one if they are fragile). Cheaper than buying the automated ones with the page turning robot arm. I have done it at home for my books since the mid-00s.

Do you purchase every book two times, or do you have a home devoid of physical books that you enjoy? Genuinely curious

Not OP, but I do similar. I only de-spine a book if it is in poor condition (and I have a nice copy). Some of the material I scan is stapled instead of glued and I can remove the staples, scan and replace with new (not rusted!) staples to return the book to its original form.

I have books but would have 5x as many if I could not capture them digitally. (When I go to move or go through a purge, some of the books I scanned do go to a used book store.)

So I scan in part to keep my physical book-footprint smaller, but also my scans all get cleaned up and uploaded to archive.org. Mainly I scan young-adult science books from the 50's and 60's (since they were so influential and have all but disappeared except on eBay and the like).

Re: AI companies are shredding rare books

#372

Earlier quoted context omitted.

I worked in an academic library that rebound pretty much every book they acquired. One of the biggest in the world, too.

Nowadays, academic libraries are moving their books to archival storage. New publications are electronic only. You need to be a formal member of the university community to access it, unlike in times past when any member of the public could stroll the aisles of physical books and journals.

Most academic libraries have large archival collections anyway, and the extent to which they de-emphasize their patron-facing stacks varies dramatically between institutions. Same with who can access them. You just can’t generalize like that.

Re: AI companies are shredding rare books

#373

I've limited sympathy for the publishers. It pisses me off to reflect that they can sit on works until copyright expires, keeping them out of print. There's no real need for any of these so-called rare books to be rare while they're under copyright. And related to this, the books that are in print are mostly only in print in the shittiest way. I often see well-made books from the 17th or 18th centuries which are stil…

I have even less sympathy for IP stealing LLM operators

It's not copying (and hence not stealing) if you destroy the original book. That is where copyright law has brought us

Re: AI companies are shredding rare books

#374
post #140

Earlier quoted context omitted.

From the court decision: "IA maintains that it delivers each Work “only to one already entitled to view [it]”―i.e., the one person who would be entitled to check out the physical copy of each Work. But this characterization confuses IA’s practices with traditional library lending of print books. IA does not perform the traditional functions of a library; it prepares derivatives of Publishers’ Works and delivers those…

Derivative works? That seems to translate as "because the books are digitized, it's strange and new and we can't allow it".

[dead]

Re: AI companies are shredding rare books

#375
post #152

Earlier quoted context omitted.

Ladies and gentlemen, the Digital Millennium Copyright Act (which is terrible, but I would be delighted if they breached it and got thoroughly spanked)

so if I base64 encode my blog, have some Javascript that 'validates' an authorized viewer and then decodes the base64 into HTML which is added to the DOM does that constitute DRM ?

Yes.

Re: AI companies are shredding rare books

#376

I've limited sympathy for the publishers. It pisses me off to reflect that they can sit on works until copyright expires, keeping them out of print. There's no real need for any of these so-called rare books to be rare while they're under copyright. And related to this, the books that are in print are mostly only in print in the shittiest way. I often see well-made books from the 17th or 18th centuries which are stil…

Books in the 17th and 18th century often didnt get bound by the publisher. They were done by independent binders for _custom_ orders. You’d see whole libraries with the owners binding/cover standards rather than per book. Books in that time were _luxury_ goods. Most people could not afford them. One of the ways that was changed was to introduce cheap, mass produced bindings that were lower quality than the bespoke ar…

The Czech National libraries do that for newspapers, magazines and other periodicals - they bind the copies they get automatically for preservation to big books, so they can be better stored in their archives.

Apparently it is getting harder to find people who can do that as most schools no longer have book binding as a course you can study.

Re: AI companies are shredding rare books

#377
post #152

Earlier quoted context omitted.

Ladies and gentlemen, the Digital Millennium Copyright Act (which is terrible, but I would be delighted if they breached it and got thoroughly spanked)

so if I base64 encode my blog, have some Javascript that 'validates' an authorized viewer and then decodes the base64 into HTML which is added to the DOM does that constitute DRM ?

[dead]

Re: AI companies are shredding rare books

#378
post #162
post #97

Earlier quoted context omitted.

Not in a traditional sense, but obviously on the surface, they're using the info to regurgitate in some fashion and serve back. The point is that even under the best intentions, hallucinations occur. Then there's the fact that most models have an ideological bias programmed into them. The only expectation I have is for companies or anybody else to not destroy rare books. Is that such a tall order?

> they're using the info to regurgitate in some fashion and serve back. Sure, in the same sense that they regurgitate any other text they consume. LLMs by definition do not have the full training dataset available, though. It’s far larger than the resulting model. So they can’t reliably reproduce full text without an external source (or if it’s in the training data repeatedly). ChatGPT actually refused to give me a b…

> It’s far larger than the resulting model.

Is it? How many different books are we talking about, and how much information is that, after conversion to text and lossless compression? Images, maybe, but text?

Re: AI companies are shredding rare books

#379

Earlier quoted context omitted.

I worked in an academic library that rebound pretty much every book they acquired. One of the biggest in the world, too.

My wife worked for a company that specialized in rebinding paperbacks for schools and libraries. It makes more economic sense to do that for niche use cases rather than make all print runs more expensive.

Yeah it depends on the library and collection. If it’s a general-purpose library, probably a pretty limited portion of the collection. If it’s a grad school library with a low-turnover collection, like the one I worked in, that’s totally different. The college I worked for had dozens of libraries and their use cases were all pretty different.

Re: AI companies are shredding rare books

#380

Earlier quoted context omitted.

Paper doesn't easily survive for thousands of years. You know that pleasant used bookstore smell? It's paper slowly decomposing.

I'd be willing to bet that there's a lot more readable paper from thousands of years ago than there will be readable digital data from today in thousands of years. Digital information has to be actively preserved in every case, whereas paper can be passively preserved in some cases.

Texts from thousands of years ago survive only because they were repeatedly copied. There are some extreme examples like the Dead Sea Scrolls but they are the exception.

Modern acid-free paper might last 1000 years; 500 is more typical. Acid paper breaks down in less than a century.

Parchment was so expensive it was often scraped and reused; old texts can sometimes be recovered after being overwritten (palimpsests).

Post reply on HN