Live data from Hacker News

AI companies are shredding rare books

twitter.com

151–160 of 559 posts

Re: AI companies are shredding rare books

#151

I've limited sympathy for the publishers. It pisses me off to reflect that they can sit on works until copyright expires, keeping them out of print. There's no real need for any of these so-called rare books to be rare while they're under copyright. And related to this, the books that are in print are mostly only in print in the shittiest way. I often see well-made books from the 17th or 18th centuries which are stil…

> I often see well-made books from the 17th or 18th centuries which are still in good nick.

Those books predate the development of wood pulp paper. It isn't the publisher's fault they can't economically print on rag paper anymore.

Re: AI companies are shredding rare books

#152
post #122

Earlier quoted context omitted.

This argument doesn't make any sense. All manner of AI companies just ingest whatever random text they can find on the internet to train their data, including copyrighted publications. Why would DRM on a digital copy of a book matter?

There might be special rules around DRM that go beyond normal copyright?

Ladies and gentlemen, the Digital Millennium Copyright Act

(which is terrible, but I would be delighted if they breached it and got thoroughly spanked)

Re: AI companies are shredding rare books

#154

We reprint old books after checking out copyrights (for all books, this means pre-1930, but for some (I'd actually say most) it also means ones published up to 1964 and 1973, depending on how the rightsholders did (or didn't) do the renewals). We use a special guillotine type cutter to cut off the binding and then store the pages in a sealed plastic bag which goes in the archives; they're stored there indefinitely in…

Some source or citation or context is needed here, is this the work of a 2 person no profit or a trillion valued pre IPO company?

Re: AI companies are shredding rare books

#155
post #3

I see mentions of Bradbury's "Fahrenheit 451" in that thread but what this really seems to be mostly like is Vernor Vinge's "shred and scan" factory in his novel "Rainbows End".

Well it's close to the author's intention for Fahrenheit 451, but just not what everyone wants it to mean.

Re: AI companies are shredding rare books

#156

We reprint old books after checking out copyrights (for all books, this means pre-1930, but for some (I'd actually say most) it also means ones published up to 1964 and 1973, depending on how the rightsholders did (or didn't) do the renewals). We use a special guillotine type cutter to cut off the binding and then store the pages in a sealed plastic bag which goes in the archives; they're stored there indefinitely in…

Why cut off the spines? Isn't that how you end up getting unattributed 'dead sea scrolls'?

Speed and cost. You can run the book through a typical sheet-fed scanner instead of using a contraption like this: https://linearbookscanner.org/

As for consumer-grade solutions, look for the Fujitsu SV600.

Re: AI companies are shredding rare books

#157
Reminds me of Blood Meridian where The Judge meticulously sketches the rock glyphs that he comes across, and then destroys the original.

Now that I think about it, The Judge is an apt metaphor for AI : "Whatever in creation exists without my knowledge exists without my consent."

Re: AI companies are shredding rare books

#158

I've limited sympathy for the publishers. It pisses me off to reflect that they can sit on works until copyright expires, keeping them out of print. There's no real need for any of these so-called rare books to be rare while they're under copyright. And related to this, the books that are in print are mostly only in print in the shittiest way. I often see well-made books from the 17th or 18th centuries which are stil…

I have even less sympathy for IP stealing LLM operators

Re: AI companies are shredding rare books

#159

And this is the beginning of the end for content creators. Why would I spend hours creating original content if Google can extract it and present the answer directly in an AI Overview? What is the incentive to keep doing the work? If creators stop producing high-quality original material, the information we get over the next few years will increasingly be based on recycled, low-quality garbage.

Take a look at most successful journalism today. It's behind a paywall. You get paid by the people who are interested and value your work.

Can Google steal it and present it in an AI overview? Well kinda. Today Google is doing a trick - they're saying "You can refuse to consent to being fed into the slop machine, but if you do we won't crawl you for Google so you'll get no search traffic. But you're not going to get search traffic anyway! So you might as well opt out of being fed into the slop machine. And companies are starting to do that [1]

It's really interesting, because essentially what it means is Google is turning into a walled garden, but there's nothing growing inside it so they have to continually import new plants to live in their walled garden and they're going to have to pay to do that. So soon Google will be paying news sites for the right to plumb their feed into the slop machine.

[1]: https://www.wsj.com/business/media/google-search-publishers-...

Post reply on HN