Live data from Hacker News

AI companies destroy physical books – let's scan rare books before it's too late

annas-archive.gl

241–250 of 961 posts

Re: AI companies destroy physical books – let's scan rare books before it's too late

#241
BOYCOTT.

It's that simple, these corps have once again broken the social contract, you must not reward them. OpenAI and Anthropic especially, both owned by schizo sociopathic elites. Just use Chinese open models on 3rd party providers or more ethical companies.

This is literally the only power you have outside of Luigi, you're not going to fix anything with a letter writing campaign. We are entering a fight for survival so you really need to step up your game and stop letting elites run over you.

Today it's just books and manipulating society, tomorrow they will track and punish your behaviour and the control will only get worse. These people are pure fucking evil and we need to start acting like it while we literally still have the freedom and privacy to organise.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#242

Earlier quoted context omitted.

Quite possibly not many, and no copy held in any form by the copyright owner either. Say a few hundred copies of some obscure book from 40 years ago. They probably won't be erased from the face of the earth by the judicious and proportionate actions of, of a few, AI companies? Hmm.

The hypothetical "heroic figure goes and buys last copy of a 1962 guide to Ford cars to carefully maintain it in an appropriately climate controlled library" is vanishingly unlikely. A ten or a hundred or a thousand times to one, it just goes to the trash. At least here it gets scanned by the AI company.

In my opinion this is one of the reasons why libraries should accept any book, even if all they do is examine it and throw it in the trash. This way they would have a chance at finding any treasures that could be regularly dumped in that way.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#243
post #40

Earlier quoted context omitted.

My personal wastebook at home is so rare, it's unique. That doesn't mean it needs preservation.

https://en.wikipedia.org/wiki/Anecdotal_evidence

You are giving me too much credit: it's made up evidence. It's an illustration that rarity doesn't equal value, and doesn't depend on whether I actually own a wastebook or ten.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#244
post #2

I dislike these AI companies but let's be clear here: the copyright holders are the ones locking these books up. If they don't want to print more copies, then they could release the copyright on them. Instead, they enforce the copyright and force AI companies to shred books they want to ingest. edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than…

They also only scan books that are easily and cheaply available, which means they are either not rare or have no significance.

Books that are rare of have historic significance will surely be in museums or libraries and not going away for pennies.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#245
post #34

The very first paragraph is fascinating: "Several AI companies are acquiring large quantities of secondhand books through intermediaries, scanning and destroying them, all to obtain training data “untouched by machines” from before 2022." Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time? Also, how useful old books really are for AI training besides helping AI acquire…

As a rule: all high quality text is useful.

There's no "2022 split", and the "untouched by machines" bit came from the marketing blurb of a company offering book scanning services - not the AI labs themselves.

At the AI lab level: the book scanning seems to be driven by copyright concerns, not data contamination concerns. There was a concern about AI contamination, but there's no measurable performance loss from ingesting post-2022 data with minimal filtration, and some tests attribute small but persistent performance gains to post-2022 AI contamination. It's unclear where exactly do those gains come from.

Why is all high quality text useful? The "inverse problem" framing is that all text reflects the thinking behind it, somewhat, and by learning to reproduce it, LLMs implicitly learn to reproduce some of the thought process too. They don't just memorize the dry factual knowledge, but also learn how that knowledge fits together, and how to reason about that knowledge - both in the specific case and in general. And that "in general" then surfaces in an LLM's ability to generalize. Which is very desirable.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#246
post #243

Earlier quoted context omitted.

https://en.wikipedia.org/wiki/Anecdotal_evidence

You are giving me too much credit: it's made up evidence. It's an illustration that rarity doesn't equal value, and doesn't depend on whether I actually own a wastebook or ten.

Your point is understood but the crux of the issue is a bit similar to capital punishment, the argument is that risk of losing even one innocent person or useful book isn't worth taking; specially considering the value created for society as part of such risky undertaking, whatever it is exercising capital punishment or scanning and destroying books.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#247

Earlier quoted context omitted.

If that's true, then why did they pirate so many books?

Is your complaint that they follow copyright law, or that they don't follow copyright law?

A company with no respect for copyright law can't pretend they respect it when it is convenient for them; it's disingenuous, and shows that there is clearly another explanation.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#248
post #2

I dislike these AI companies but let's be clear here: the copyright holders are the ones locking these books up. If they don't want to print more copies, then they could release the copyright on them. Instead, they enforce the copyright and force AI companies to shred books they want to ingest. edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than…

It's also the regulation, where most systems still look at "one pirate copy" = "one sale of lost profits", especially when pirates end up in court. If the book (or game or whatever) is not sold anymore in any way where you could give the copyright holder money in an easy accessible way (eg. buy it on amazon, or a local bookstore), they shouldn't be able to claim losses from piracy, since they clearly don't want your money.

On the other hand, there are grey zones here, the lord of the rings books (still copyrighted and easily obtained pretty much everywhere) have been translated into my language many decades ago, and many of us read and liked those translations, but when the movies came out, a new translator did a new translation, where they changed a lot of things, including the last names of bilbo and frodo (Bogataj->Bisagin) and the Shire (Grofija->Šajerska), and the old version is sadly available only in paper form on second hand markets. On one hand, copying that if you only want this specific version would not cause a lost sale, on the other, you can get new translations (or english originals) pretty much everywhere.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#249
Do you mean rare as in old? I doubt they are destroying old books because most of them are out of copyright and probably available already as text. I imagine this applies to in copyright works and I do NOT condone it but the way this is told it sounds like they are raiding old libraries to destroy first editions of Cervantes.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#250
post #188

Earlier quoted context omitted.

Or not being able to. Regional licensing, missing and shuffling content. To watch the world cup I had to spin up a VM in Brazil to watch it with Portuguese narration because the free transmissions are region locked. I would gladly pay 5 bucks for it if it was possible otherwise and avoid the hassle.

There was no way in your country to pay and watch it? FIFA will have sold the tv rights there to someone, surely. In which case your complaint is what, that it was expensive?

Free on both, but not with the narration in my native language.

On Brazil the world cup was being transmitted on youtube. in NL only on traditional TV channels or Online for the same channels (all for free but in Dutch).

And literally as I write this I receive an email saying that my youtube premium was raised from 33 to 38 EURO. So there we have piracy getting juicier and juicier.

Post reply on HN