Live data from Hacker News

AI companies are shredding rare books

twitter.com

31–40 of 559 posts

Re: AI companies are shredding rare books

#31
post #16

That's why archive.org should have never been sued for lending books they had physical copy of. This is the result. Publishers should be more careful what they wish for.

Publishers don't care if rare books get shredded?

And, regrettably, The Archive lent books regardless of physical possession.

Publishers had accepted the prior arrangement before The Archive decided to push it, if not explicitly then implicitly by not suing.

I'm a believer in The Archive's mission, and I wish they had treated the goodwill they'd accumulated as something worth preserving and not a currency to be spent.

It has been stated by many before me: lending books should have been handled by a separate entity, especially when they removed the physical backing requirement.

Re: AI companies are shredding rare books

#32
This kind of reads like a blood libel. My guess is they're buying all those books that University libraries are throwing away these days (to see other HM threads for that), kinda sad but probably aren't "rare books" in the way people are thinking

Re: AI companies are shredding rare books

#37

The publishers sued AI companies for training on shadow library data, hoping to negotiate content deals for big $$$ down the line. Instead, they got analog hole'd. Turns out that buying an old book for $5 and destructively scanning it for $25 is way cheaper than paying extortion fees to the copyright-mongers. What I don't buy is it being "rare, precious books". First, they're not after ancient texts - they're after t…

I see.. so the various AI companies are in the right on this?

I think they are in the legal sense of right, and I think they only discarded the remains of these dissected books because previous rulings (e.g. archive.org's lending practices of digital copies of books they physically owned) gave rise to a situation where destruction bore less legal risk. As for the moral case, I don't have much to say on that, we all have our own lines in that sand.

Re: AI companies are shredding rare books

#38
post #16

That's why archive.org should have never been sued for lending books they had physical copy of. This is the result. Publishers should be more careful what they wish for.

The archive.org story was more nuanced than that. If I recall correctly the full story was that they used to lend digital versions of books they physically bought and scanned with DRM to enforce a sort of one to one at a time restriction.

But during covid archive.org decided to just remove the limit and lend unlimited copies concurrently which started the debacle with the publishers.

Re: AI companies are shredding rare books

#39
post #17

I think people are, on the whole, too precious about old things. In the case of books produced after major commercial printing began, I don't believe it is the paper that imbues the book with historical value. Indeed, I think there's a high chance that this process increases preservation of the most relevant part of the media - the actual content! There are tons of old rolls of film slowly rotting away in warehouses…

> For now these books are in corpuses of training data, but eventually I trust they will make their way to the rest of us. What makes you think they will? What would be the incentives for these companies to do so?

Well, you probably weren't going to go and find any rare, non-digitised books to physically go and read (unless you were going to, in which case, rock on), so we can start by benchmarking relative probability there.

1. At some level of critical information-withholding mass, a leak or disclosure similar to SciHub is inevitable because of the commonly held opposition to hiding knowledge.

2. Availability via Google Books or similar.

3. Availability via AI model reference.

4. Failing any of the above, better AI models that are more capable of doing more things, at the expense of books that were likely to go unread (revealed preference, rare books are often rare for a reason). This will obviously be a nonstarter if you don't want this to happen, but I think it would be good for the world if it did.

I think category of old books that were going to be read or otherwise become important parts of human knowledge that have not yet been digitised and now will never become so because they are instead being shredded and will never make their way into the light because of AI company data hoarding is a small category.

Re: AI companies are shredding rare books

#40

The publishers sued AI companies for training on shadow library data, hoping to negotiate content deals for big $$$ down the line. Instead, they got analog hole'd. Turns out that buying an old book for $5 and destructively scanning it for $25 is way cheaper than paying extortion fees to the copyright-mongers. What I don't buy is it being "rare, precious books". First, they're not after ancient texts - they're after t…

I see.. so the various AI companies are in the right on this?

[deleted]
Post reply on HN