Live data from Hacker News

AI companies are shredding rare books

twitter.com

101–110 of 559 posts

Re: AI companies are shredding rare books

#101
post #73

Librarians were already doing this at scale in a process euphemistically called "weeding": https://www.ala.org/tools/challengesupport/selectionpolicyto...

This is a bizarre comparison.

Weeding is the natural process of disposing of less-demand books. Like the rest of us, libraries operate in finite space, so if they want new books, they have to remove ones their users aren't using. Most libraries will try to sell books before disposing of them in any destructive way.

What similarities do you see here?

Re: AI companies are shredding rare books

#102
People would have at least somewhat less of a problem with this if they also put up an archive of PDFs of all these rare books if they are out of copyright.

But that would help competitors with training data, which I assume is why they don’t do this.

Re: AI companies are shredding rare books

#103

Earlier quoted context omitted.

> Barrett's Traditional Fairy Tales (2021) How is a book from 2021 considered rare in this context? There's almost certainly a digital copy of it in existence prior to Anthropic purchasing a print edition.

Niche text. It's not impossible that there was only ever under a thousand of them printed and released into circulation. A digital copy would exist somewhere, of course. But for us, that only matters if we can buy or download it. And for AI companies, that only matters if they can get a digital copy DRM-free and licensed permissively enough.

This argument doesn't make any sense. All manner of AI companies just ingest whatever random text they can find on the internet to train their data, including copyrighted publications. Why would DRM on a digital copy of a book matter?

Re: AI companies are shredding rare books

#104
post #63

This is the opposite of a book burning. These books which only a few would ever know the names of, let alone find, let alone read, are being digitised so they can be found in electronic searches.

And knowledge from these books can be "imprinted" into a model to be used by much more people or even survive this planet once sent into space in a probe.

Re: AI companies are shredding rare books

#105
post #85
post #73

Librarians were already doing this at scale in a process euphemistically called "weeding": https://www.ala.org/tools/challengesupport/selectionpolicyto...

This is really dishonest framing, unless you really, honestly can't tell a difference between pulping a mass market paperback romance novel that there's 3 million of in circulation, and shredding an 18th century botanical text that there's only 2 copies of in existence.

You call of dishonest framing, but you're begging the question twice.

Once that AI companies are really shredding 200 year old rare books, and once that libraries are only weeding mass market pulp fiction.

Re: AI companies are shredding rare books

#106
post #71

I don't see any proof of shredding here. Most book scanners I'm aware of are from Google's scanning days, and those had cameras plus page turning. If we use 'shredding' to mean A book is laid flat, its cover is removed, and then a paper cutter cuts through the binding to create a flat stack of sheets, which are then fed to a sheet feeder, then I could maybe imagine this is better than a page-turning scanner. But, she…

The idea here is that the frontier labs are no longer willing to risk using the likes of Anna’s Archive, because they have already been held legally liable for that in the past.

Are you referring to the META lawsuit? I think the current landscape is not bad for the frontier labs -- it's settled law that it's legal to 'read' and ingest this data. Anthropic went ahead and just settled a licensing deal for content. The open issue in that META suit is whether or not any distribution happened, as I understand it. I'm certain they all have full backups of the archive somewhere in the org.

Re: AI companies are shredding rare books

#108

Earlier quoted context omitted.

That’s digital though, so it requires the continued survival of readers for the data that is stored. The best thing you could for the long term is probably to buy a few hundred physical books to keep in a bookcase in your home.

Speaking as someone with dozens of bookshelves and tens of thousands of books... I kind of prefer the idea that continued survival means getting a bunch of 14TB drives and, you know, hosting certain files obtained from certain places. The reality is that most people's book collections are simply going into dumpsters, speaking as one of the people who go and try to buy these book collections at estate sales. (We can't…

Yeah agreed. I have on my long-term project list a hardware 'oracle' that would have everything and a local good model as a librarian/assistant and be solar powered in a pinch.

Re: AI companies are shredding rare books

#109
post #16

That's why archive.org should have never been sued for lending books they had physical copy of. This is the result. Publishers should be more careful what they wish for.

The archive.org story was more nuanced than that. If I recall correctly the full story was that they used to lend digital versions of books they physically bought and scanned with DRM to enforce a sort of one to one at a time restriction. But during covid archive.org decided to just remove the limit and lend unlimited copies concurrently which started the debacle with the publishers.

> But during covid archive.org decided to just remove the limit and lend unlimited copies concurrently

IIRC, this was 100% it. Lending one digital version of one physical asset was likely already a violation copyright. Lending UNLIMITED digital versions of one physical copy was DEFINITELY a blatant violation of copyright.

Re: AI companies are shredding rare books

#110

Earlier quoted context omitted.

Scanning books by taking them apart into singular pages and scanning those pages is faster and cheaper. AI training is a numbers game, so they want faster and cheaper. What happens to the pages after? No one needs them anymore, so they get mulched and recycled. That would be the dominant scanning method even if copyright wasn't a thing. But then again - if copyright wasn't a thing, there would be much less need to sc…

> if copyright wasn't a thing, there would be much less need to scan any physical media. Because there’d be much less content created in any media to capture in the first place.

Empirically, probably not. We had lots and lots of content before copyright, and people seem to produce lots of content even in jurisdictions with weaker copyright.
Post reply on HN