Live data from Hacker News

AI companies are shredding rare books

twitter.com

281–290 of 559 posts

Re: AI companies are shredding rare books

#281

We reprint old books after checking out copyrights (for all books, this means pre-1930, but for some (I'd actually say most) it also means ones published up to 1964 and 1973, depending on how the rightsholders did (or didn't) do the renewals). We use a special guillotine type cutter to cut off the binding and then store the pages in a sealed plastic bag which goes in the archives; they're stored there indefinitely in…

> And no, no AI company has ever come to us and asked to run training on all of our scanned copies

The value of most very old books for AI training is very low. You don’t really want your AI training data to start biasing toward outdated writing styles. Most of the valuable knowledge has been covered again in modern texts in more depth and detail.

There is interesting value in old texts and it’s important to have them archived. It’s less valuable for stirring into the giant pot of AI training data, though.

Re: AI companies are shredding rare books

#282

I've limited sympathy for the publishers. It pisses me off to reflect that they can sit on works until copyright expires, keeping them out of print. There's no real need for any of these so-called rare books to be rare while they're under copyright. And related to this, the books that are in print are mostly only in print in the shittiest way. I often see well-made books from the 17th or 18th centuries which are stil…

I'd go for the less complex version: just make copyright last for like 20 years at most.

Re: AI companies are shredding rare books

#284

Earlier quoted context omitted.

A friend proposed that copyright should just die with the author and / or their spouse and I'm left agreeing. I want books and music to be less strict on copyright. Some of my favorite YouTube channels break down music and songs, and go as far as recreating beats / tracks from famous hip hop songs, but someone at a record label company dings every one of their videos, they can barely sample a few seconds, its VERY CL…

This would create incentive to kill people

So does life insurance.. but we still have that.

Re: AI companies are shredding rare books

#285

Earlier quoted context omitted.

Think about that last point for a moment. Our “rights to read” are diminished significantly with digital works as compared to printed works. Right of resale. Right to lend. In the end, digital publishing just isn’t right and will lead to massive gap in our historical records. They require active curation and cannot be preserved simply by resting on a dusty shelf.

Every innovation since the microprocessor isn't worth saving in the grand scheme of things. When today's algae evolve enough into tomorrow's sentient creatures, they're really only going to need up to the industrial revolution and should probably stop right before that.

I'm personally a fan of more than 50% of children surviving past the age of 6, something that didn't happen until the 20th century.

Re: AI companies are shredding rare books

#286
post #204

Earlier quoted context omitted.

No, courts so far in the jurisdictions which have heard such cases, have ruled it's fair use. There is plenty of ongoing litigation in many jurisdictions, so it's way too early to just decree "it's been determined". It likely won't be for years to come.

Why did Anthropic settle with authors for 1.5 billion then? Surely their lawyers must have decided there's a pretty good chance of judges ultimately deciding that it is copyright infringement?

You could also ask why the other side agreed to that settlement. It's not a one-way street.

Re: AI companies are shredding rare books

#287
post #162

Earlier quoted context omitted.

> they're using the info to regurgitate in some fashion and serve back. Sure, in the same sense that they regurgitate any other text they consume. LLMs by definition do not have the full training dataset available, though. It’s far larger than the resulting model. So they can’t reliably reproduce full text without an external source (or if it’s in the training data repeatedly). ChatGPT actually refused to give me a b…

> LLMs by definition do not have the full training dataset available, though. That makes it even worse, then. This proves the original point. > The idea that people or corporations should hold onto books forever because of a cultural “ick” about throwing out books is a bit ridiculous. If we're building black and white straw man arguments, then sure, let's not archive anything.

> That makes it even worse, then. This proves the original point.

I don’t know what the “original point” is here, but these AI companies are not providing “book excerpt services” and do not claim to. ChatGPT at least will refuse to provide detailed book excerpts (I hit a week or two ago myself).

> If we're building black and white straw man arguments, then sure, let's not archive anything.

It seems like you are the one creating the straw man. Do you have evidence that these companies are shredding actually rare books? The only cited concrete examples (in this thread anyway) are all rather boring. I seriously doubt they are shredding 200 year old books because why would they?

Re: AI companies are shredding rare books

#289
post #257
post #220

Earlier quoted context omitted.

[flagged]

I genuinely want to know what exact point you are trying to make, I'm not a mind reader.

This entity is destroying books and make them only accessible via the "intelligence" exposed by their LLMs. So I ll give you two options...

* their values are aligned with the best interests of humanity

* their values are not aligned with the best interests of humanity.

Now go ahead and read my mind!

Post reply on HN