Live data from Hacker News

AI companies are shredding rare books

twitter.com

501–510 of 559 posts

Re: AI companies are shredding rare books

#501

Earlier quoted context omitted.

It's possible. My only experience is with vintage newspapers that were allowed to dry out, and then when trying to open them just shatter like a carbonized volcano scroll.

It's truly humbling how much knowledge professionals from different fields have, and how easy it is to fall into a Dunning–Kruger effect trap where a bit of knowledge — extremely extrapolated — would've led me down the path of hubris and guessing so many wrong answers. Thanks all.

One of the best things about this site is how many people there are on here who are way smarter than me, with much deeper domain knowledge.

Re: AI companies are shredding rare books

#502

I've limited sympathy for the publishers. It pisses me off to reflect that they can sit on works until copyright expires, keeping them out of print. There's no real need for any of these so-called rare books to be rare while they're under copyright. And related to this, the books that are in print are mostly only in print in the shittiest way. I often see well-made books from the 17th or 18th centuries which are stil…

A friend proposed that copyright should just die with the author and / or their spouse and I'm left agreeing. I want books and music to be less strict on copyright. Some of my favorite YouTube channels break down music and songs, and go as far as recreating beats / tracks from famous hip hop songs, but someone at a record label company dings every one of their videos, they can barely sample a few seconds, its VERY CL…

I thought about this and what they are doing is writing the music industry out of the minds of the next generation who mostly uses Youtube etc. for entertainment.

Most Youtube videos do not contain any copyrighted music as it takes too much revenue as you state.

Again, its short term profit at the cost of long-term gain.

Re: AI companies are shredding rare books

#503

Earlier quoted context omitted.

It's truly humbling how much knowledge professionals from different fields have, and how easy it is to fall into a Dunning–Kruger effect trap where a bit of knowledge — extremely extrapolated — would've led me down the path of hubris and guessing so many wrong answers. Thanks all.

One of the best things about this site is how many people there are on here who are way smarter than me, with much deeper domain knowledge.

Or we just reflexively look stuff up.

Re: AI companies are shredding rare books

#504

Earlier quoted context omitted.

The archive.org story was more nuanced than that. If I recall correctly the full story was that they used to lend digital versions of books they physically bought and scanned with DRM to enforce a sort of one to one at a time restriction. But during covid archive.org decided to just remove the limit and lend unlimited copies concurrently which started the debacle with the publishers.

> But during covid archive.org decided to just remove the limit and lend unlimited copies concurrently IIRC, this was 100% it. Lending one digital version of one physical asset was likely already a violation copyright. Lending UNLIMITED digital versions of one physical copy was DEFINITELY a blatant violation of copyright.

sure, but there's no way you believe that that's not the way they should've been doing it all along, right?

Re: AI companies are shredding rare books

#505

Earlier quoted context omitted.

Either that or some exponentially increasing tax so that Disney can keep their vault. (I'm perfectly fine with them keeping it if they pay some proper taxes.)

the important thing is that copyright serves the public good, I do not think it currently does, at least not as well as it could

The whole purpose is to promote the progress of science and the useful arts.

The implementation is no longer in alignment with that constitutional requirement.

Re: AI companies are shredding rare books

#506

We reprint old books after checking out copyrights (for all books, this means pre-1930, but for some (I'd actually say most) it also means ones published up to 1964 and 1973, depending on how the rightsholders did (or didn't) do the renewals). We use a special guillotine type cutter to cut off the binding and then store the pages in a sealed plastic bag which goes in the archives; they're stored there indefinitely in…

> And no, no AI company has ever come to us and asked to run training on all of our scanned copies The value of most very old books for AI training is very low. You don’t really want your AI training data to start biasing toward outdated writing styles. Most of the valuable knowledge has been covered again in modern texts in more depth and detail. There is interesting value in old texts and it’s important to have the…

It could be translated to contemporary prose before training.

Re: AI companies are shredding rare books

#508

Earlier quoted context omitted.

breaking DRM is so easy my generation was doing it as kids, there is no technical obstacle there.

It's a manual process. More importantly, it's also explicitly illegal. Destructive format-shifting is not. Thank copyright laws.

Why would it be a manual pass? There's only so many DRM schemes, and you can even use your previous generation AI to help with the breaking.

> More importantly, it's also explicitly illegal.

Agreed!

Re: AI companies are shredding rare books

#509

I've limited sympathy for the publishers. It pisses me off to reflect that they can sit on works until copyright expires, keeping them out of print. There's no real need for any of these so-called rare books to be rare while they're under copyright. And related to this, the books that are in print are mostly only in print in the shittiest way. I often see well-made books from the 17th or 18th centuries which are stil…

This has little to do with copyright, chances are many of these books were written before that. It’s much more about preserving knowledge. AI labs are a bit like the Borg. They eat up knowledge and so far the average output is often questionable.

Re: AI companies are shredding rare books

#510

Earlier quoted context omitted.

You don't need to shred books to scan them. They make book scanners that will "rip" a fully-bound book no-problem, and even correct for the curvature of the page and binding to give you an equivalent image. In fact, the Internet Archive specifically built nondestructive book scanners[0] for exactly the purpose of which AI companies are now shredding books. The smart / savvy thing to do would be to buy those machines…

Again, the entire point I'm making is that the choice is dictated by copyright law . They can't legally scan books non-destructively. They can legally format-shift them , i.e. scan them destructively. So this is what they're doing. Your comment seems to also be regurgitating common misconceptions (to put it charitably) about AI and web scraping.

There is no tenet of copyright law that requires format-shifting be destructive. The term "format-shifting" almost always refers to a non-destructive process; i.e. when you "rip" a CD you are getting a 1:1 copy of the music, but the original CD still exists. If you were legally expected to destroy the disc after ripping, RIAA v. Diamond would have gone a different way and MP3 players would have been illegal.

For AI training specifically, the only standing caselaw is the Anthropic lawsuit. And in that lawsuit, the only thing that was actually in the wrong was maintaining a library of pirated books. That was deemed illegal and Anthropic was ordered to delete those files. But, notably, the judge explicitly said that scanning books to train AI on them was legal, and imposed no requirement to destroy scanned books. I'm pretty sure Anthropic wouldn't even need to retain the physical copies - though there's no caselaw on that in particular, so don't cite me.

Post reply on HN