Live data from Hacker News

AI companies destroy physical books – let's scan rare books before it's too late

annas-archive.gl

411–420 of 961 posts

Re: AI companies destroy physical books – let's scan rare books before it's too late

#411
I imagine they are only buying one copy of each book, thus only significantly affecting the supply of books that were already unfathomably rare. That may still be bad, but doesn't really support the "scan every book you can get your hands on before they are gone" narrative.

That doesn't mean I'm against that narrative; I'm a big supporter of shadow libraries and scanning every unscanned book. But connecting to the LLM narrative here seems opportunistic and populistic.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#413
post #352

Earlier quoted context omitted.

They cannot scan the books then resell them or donate them under current US copyright law. It’s not clear to me that they could warehouse them if they wanted. In the recent Bartz v Anthropic case the judge ruled this destruction as legal, saying > The print original was destroyed. One replaced the other. So that there was still only one “copy” of the book. This is in compliance with the DMCA. You can make a personal…

Nowhere in this description did it require destruction of the physical book. This is being done because it's easier to scan a shucked book, and this explanation is circulating because it's easier to blame it on the law and that pesky meddling government.

So if I scan a book, sell it, and keep using the scan, is that legal? (Spoiler: That's not legal. It's a violation of IP law.)

Re: AI companies destroy physical books – let's scan rare books before it's too late

#414
post #299

I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collective responsibility. Thus, the preservation of this knowledge, its availability, and accessibility are t…

A small change in the copyright law would fix this problem. Something like: If a company is scanning material protected by copyright, it has to send a digital copy of the scanned material to Library of Congress within 5 working days.

So now the Library of Congress has to manage all these submissions whenever someone scans something? How do you even go about enforcing such a thing?

Re: AI companies destroy physical books – let's scan rare books before it's too late

#415
post #409

I have a year membership to AA now, been meaning to contribute for some time, archive.org is next on the list Much like the pushback we are seeing from the citizenry against things like Flock and AI datacenters, we can push _forward_ too by ensuring important institutions (legal or otherwise) remain funded

archive.org is an incredible blessing that I never really respected enough until the last few years. I have bookmarks going back to the 90s, and for some reason, starting in 2015 or so, sites just started disappearing. I estimate at least 20% of my bookmarks are 404 now.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#416

The piracy organizations are playing 4D chess while everyone else is playing checkers. The irony of this entire situation - AI companies being legally required to shred books due to kafkaesque copyright laws, then used as a marketing tactic by Anna's Archive - is a work of art. I support Anna's Archive, by the way. Information wants to be free.

Why are the companies legally required to shred the books? That's the most surprising part about this to me. Surely if they bought them second hand they could donate or resell after scanning. I'm wondering if the scanning machines are damaging the books.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#417

Earlier quoted context omitted.

> Show me which “rare” books they are destroying and _maybe_ I’ll care but so far the pearl-clutching over this BBC good enough for you? https://www.bbc.com/news/articles/cp3rprx2wl4o "A recent academic text published in only 100 copies, 75 of which are already in libraries, may be very rare on the market - but it is perhaps not such a great loss if one copy is destroyed," says Derek Walker, owner of Edinburgh booksh…

It would be if it made the point you think it is. All of this is more and more hand waving. 75 copies in museums? Then I think we’re good. As for the 18th century books, that’s pure speculation. It’s like a museum saying, “yes we have they prints for sale that they keep buying and destroying but wouldn’t it be a shame if someone destroyed the actual Mona Lisa?”. And lastly, if these books are so important, then don’t…

> It would be if it made the point you think it is. All of this is more and more hand waving. 75 copies in museums? Then I think we’re good.

You're repeating the seller's contextual point, as if it's a counterargument. Do I need to explain to you that what makes the lone surviving 18th century edition important is that there aren't 75 copies of it in museums?

> And lastly, if these books are so important, then don’t sell them, hold onto them, digitize them without destroying them.

It's good to see you agree any digitization of this category of book should be non-destructive.

> If these books are so rare and important, then why has no one cared until now to actually preserve them?

(Lastly for realz this time, eh?) Why has no-one cared to actually preserve the actually preserved book being sold by the bookseller... Bit of a strange question, that.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#418

You ask "Why destroy physical books?" I ask "Why save physical books?" If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere. Owning and storing physical books is not free, there is a real cost. Even owning and storing the scans is not free, especially when IP rights and challenges get involved, since a scanned book with no distribut…

> If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere. The Gutenberg bible is rare, its content is not rare, and some would argue that the content itself is not valuable, and yet the Gutenberg bible is valuable. Same can be said about many books.

Ask: valuable to whom and why?

- Reader: narrative

- Collector: scarcity of the physical artifact

- AI Company: language samples (quantity, variety), facts

Re: AI companies destroy physical books – let's scan rare books before it's too late

#419

The main question is why aren't they leaking it to AA themselves? Trying to keep an edge with their training sets? Isn't it ridiculous, considering the sheer size of them and statistical insignificance of the set differences?

> The main question is why aren't they leaking it to AA themselves?

How is this a question at all? They’re scanning books because the courts determined that it’s the only way to use that data. They are forbidden from using digital copies found on places like Anna’s Archive. They must acquire and scan the book.

They cannot redistribute the book. The Internet Archive tried that and the courts shut it down. You cannot scan a book and share it without violating copyright law.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#420
Sorry for the copy-paste from elsewhere. I'm probably doxxing my online accounts with this. 100% my words tho and 100% I stand by this.

Really can't wait for this outrage cycle to finally die. There's a lot AI companies are doing wrong. The wording of Project Panama is really some villain-type shit. But if you stop and think about it, this is really not concerning. Like, at all.

0. In the Anglosphere, the British Library (UK and Ireland) and the US Library of Congress (USA) preserves all books published in their territories. So if Anthropic is pulling a 1984/Fahrenheit 451, "gating access to information", they are really doing a ridiculously bad job at it. I'm pretty certain similar depositories exist for other jurisdictions.

1. Rare does not automatically mean it has some obscure knowledge nor does it mean culturally/historically significant. The oldest book the news outlets could mention is a 1970s manual on soil mechanics (Guardian link in your sources). How "obscure" is the knowledge in there, you reckon? How culturally significant is that book? Odds are, a good chunk of that book is outdated knowledge at this point, useless for anything practical other than knowing what people in the 70s thought about soil mechanics. And for the latter, there are a bunch of other soil mechanics manuals from the 1970s lying around.

2. No, despite the admittedly cartoonishly villainous description of Project Panama, Anthropic is not out to get every single manual on soil mechanics from the 1970s out there. They just need to cover enough subject breadth in their training data set. Getting redundant copies is a waste of money. See also, point 0.

3. You argue a lot for the nostalgia and sentimentality of old books, which, okay, but that is hardly beside the point. Libraries and publishers destroy books at a greater scale than Project Panama when there's not enough demand for them. That's not an affront to history or culture just the consequence of economics (see also, point 0). Similarly, y'all outraged about old and second hand books being destroyed but odds are, they would've been trashed anyway even if Anthropic didn't get their hands on them. We've had all the time in the world for someone else to buy them to keep them from big bad Anthropic but no one did. I'm just happy for the booksellers who made some money out of this AI-infected timeline.

4. It's not like Anthropic is buying books from Dr. James Justin Sledge---books not covered by point 0 in other words---and chopping that up. The fact that one of their reported suppliers is "ISBNdb" should clue you in that they are interested in books from the ISBN era (and hence covered by point 0). Fact is, it's a terrible investment to "gate knowledge" from the really rare and old and culturally significant books. They are über expensive and for what? Even more useless is the AI trained on those books! Imagine asking ChatGPT how to cook a large squid and it answers in the style of Thomas Hobbes.

Post reply on HN