Live data from Hacker News

AI companies destroy physical books – let's scan rare books before it's too late

annas-archive.gl

291–300 of 961 posts

Re: AI companies destroy physical books – let's scan rare books before it's too late

#291
We've been seeing that headline for a few weeks now and I really don't understand the problem.

Companies are paying for the books now, great! And destroying a book is really the best way to scan it. I could destroy the books I buy myself if I wanted, they're my books and I can do whatever I want with them. A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books are printed in thousands to millions of identical copies nowadays.

Now what books exist in a single copy that destroying it would amount to losing knowledge? In any case, I would also expect such books to be very old, have very little useful content for training to start with, and be expensive enough to buy to make training on them unprofitable.

So what's the problem here exactly?

Also from the article:

> It’s outrageous is that it’s legally permissible, but ethically, it’s an extremely serious crime against humanity.

I'm all love for Anna's archive, but still, I find it rich that they now find themselves in position to make strong ethical claims.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#292
post #285

Earlier quoted context omitted.

At least there will be a copy left for us. The AI companies won't share these books in their original form.

Because it's illegal. That's the whole reason they are shredding books in the first place, because copyright law forces them to do stupid things. Google wanted to share the whole of Google Books 15 years ago, too, but they were sued to hell, so now you get a watered down search functionality.

Yes, government regulations are almost always behind commercial entities making seemingly irrational choices.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#293
post #285

Earlier quoted context omitted.

At least there will be a copy left for us. The AI companies won't share these books in their original form.

Because it's illegal. That's the whole reason they are shredding books in the first place, because copyright law forces them to do stupid things. Google wanted to share the whole of Google Books 15 years ago, too, but they were sued to hell, so now you get a watered down search functionality.

The poow AI execs being forced to commit acts of intewwectual tewwowist when all they wanted was to cynicawwy make the wowld a wowse place

Re: AI companies destroy physical books – let's scan rare books before it's too late

#294
post #291

We've been seeing that headline for a few weeks now and I really don't understand the problem. Companies are paying for the books now, great! And destroying a book is really the best way to scan it. I could destroy the books I buy myself if I wanted, they're my books and I can do whatever I want with them. A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books ar…

The problem is we don't know what we're losing, due to lack of transparency.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#295
post #291

We've been seeing that headline for a few weeks now and I really don't understand the problem. Companies are paying for the books now, great! And destroying a book is really the best way to scan it. I could destroy the books I buy myself if I wanted, they're my books and I can do whatever I want with them. A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books ar…

I think the problem is precisely that for some books there are not too many copies around like you described and if they destroy them we might eventually lose access to them directly. It might sound too extreme, but I understand the fear behind this.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#296
post #291

We've been seeing that headline for a few weeks now and I really don't understand the problem. Companies are paying for the books now, great! And destroying a book is really the best way to scan it. I could destroy the books I buy myself if I wanted, they're my books and I can do whatever I want with them. A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books ar…

The problem is that you probably do little research or read very few old books.

There are multiple instances of a book being referred within another book, while at the time the author had access to it, we might not have it today. Taking a rare book and destroying absolutely erases that link we have with the past.

You not seeing a problem with this is the core issue, it's probably why the people doing it (it's people destroying these books not aliens) just shrug and don't feel too bad doing that.

Old books are even more crucial than today's books due to how uncommon it was to have something written/printed and bound. Many unique and single copy books explain to us a ton of things about the past, sometimes for funsies and sometimes for useful findings. Destroying old books is akin to destroying the closest we got to time machines.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#297
The main question is why aren't they leaking it to AA themselves? Trying to keep an edge with their training sets? Isn't it ridiculous, considering the sheer size of them and statistical insignificance of the set differences?

Re: AI companies destroy physical books – let's scan rare books before it's too late

#298
post #285

Earlier quoted context omitted.

At least there will be a copy left for us. The AI companies won't share these books in their original form.

Because it's illegal. That's the whole reason they are shredding books in the first place, because copyright law forces them to do stupid things. Google wanted to share the whole of Google Books 15 years ago, too, but they were sued to hell, so now you get a watered down search functionality.

> because copyright law forces them to do stupid things

This is such dangerous train of thought, to give them the benefit of being forced to destroy books. Why is that exactly, and who is forcing them? You can also, you know, find another way?

Like the data centers who currently use very dirty energy acquisition methods (not all of them), are they also "forced" to do this, because they too need to make as much money as the other ones? How long would you continue this idea of others "forcing" for-profit companies to try to make more money, regardless of consequences?

Destroying books used to be an obvious dumb, stupid and shit idea, not sure how somehow a for-profit company making of a digital copy for themselves of the book before destroying it, suddenly makes it not a shit idea for the rest of humanity.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#299
I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collective responsibility. Thus, the preservation of this knowledge, its availability, and accessibility are the most important things that we must ensure continue.

From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.

Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available. As stated in my first paragraph, they are not the exclusive stewards of humanity despite them anointing themselves as such. Granted, a lot of these books might not be that useful, but still a relic of times pre-machine generated text, which makes them valuable if only for their archival value.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#300

Earlier quoted context omitted.

Lossily stuffing books into a model through a training process has so far been deemed legal. Enabling piracy by giving away digital copies is not. The internet archive tried to give away digital copies and they got sued to hell and back (which they should've seen coming from miles away). I'm sure they have some repository available somewhere. They can even sell the digital copies down the line if they're done with th…

>Lossily stuffing books into a model through a training process has so far been deemed legal Not in the EU, UK, or US. "AI" companies were already forced to settle their piracy cases, but often they get a free pass by law enforcement via regulatory capture. The problem is a book author contracted publisher does not assign legal rights of duplication to a company/individual that buys a legitimate print. It can take ov…

The piracy lawsuits have so far only deemed that obtaining books through pirating is a violation of copyright. Training the models on them and reselling model access has yet to be ruled illegal, from my understanding, even though there has been plenty of opportunity to.

Anthropic's crime wasn't stealing the contents of books and making a derivative work of it, but torrenting a shitload of books. Had they bought all the ebooks, I don't think the lawsuit would've gone anywhere.

Post reply on HN