Live data from Hacker News

AI companies destroy physical books – let's scan rare books before it's too late

annas-archive.gl

901–910 of 961 posts

Re: AI companies destroy physical books – let's scan rare books before it's too late

#901

Earlier quoted context omitted.

If they have no value why are they being acquired and scanned?

Any original text has value for the purpose of AI training. The argument is that they have no other value.

Exactly. It’s the funniest thing. These books are only being bought by AI companies. It seems pretty clear they only have value to AI companies. Otherwise all these people bemoaning the loss of rare books would be… buying them.

“How dare these companies buy rare and valuable books that nobody else values enough to buy” is a self-canceling argument.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#902

Earlier quoted context omitted.

Because destroying the book makes it possibly inaccessible permanently because we have no idea how long anthropic plans on storing the digital copy, if at all. A lot of people assume they would for future training, but you don't know that. If they at least didn't destroy the book, someone could purchase it when anthropic eventually goes belly up. Hopefully someone will at least be able to purchase their digital scan…

> we have no idea how long anthropic plans on storing the digital copy, if at all. A lot of people assume they would for future training, but you don't know that. Books are considered super high quality training data. Anthropic has no reason to get rid of this data that 1. They’ve spent a ton of money on and 2. Will remain useful indefinitely for training LLMs. > If they at least didn't destroy the book, someone coul…

Yesterday you could have purchased one of these books and read it. Good luck doing so today.

I don't know why you're so quick to assume this will all just "work out" such that the scans are ultimately accessible. Arguably the most likely scenarios are either Anthropic survives and holds them away in perpetuity or Anthropic fails and they are sold off to the highest bidder who does the same.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#903
post #870

Earlier quoted context omitted.

Why assume it is? We don't know, and the fact that the process is so opaque and would have been totally undisclosed if not for journalists doesn't exactly make me confident they are "doing things by the book" on this one.

Because you usually find out if someone is doing something before you get mad about it. No?

You should already be upset. A company with far greater resources than you has potentially prevented you from accessing a portion of the world's information. Because they haven't been transparent, you don't know. You should react to the facts, which are that a company has taken moves, largely surreptitiously to prevent you from accessing intimation. There is zero reason to assume good will in this scenario. You should demand more transparency and more insight into operations, period. Unless you have a vested interest in the company.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#904

Earlier quoted context omitted.

How would film be any different in any capacity whatsoever? Believe it or not, what matters here is the message and access to the message, not the medium.

The medium obviously matters. One was designed to be cheap and mass produced. Almost anything published has copies in the hundreds or thousands at least and the chances you are concerning yourself with the 'last copy' of some valuable information you can't get elsewhere is minuscle. Film most certainly was not like that.

Assuming these texts exist in large quantities seems to definitionally contradict the term *rare* book. Plenty of books have small print runs or have reached out of print status.

One could obviously reproduce films too.

Ultimately, you don't know how big N is, and that's what matters. Part of the problem is there is no transparency around this. The population of all books != the population of a specific book or set of books. Assuming N is large when you have basically zero information on specifics is foolish. These books might be just as rare as a film for which only one source exists you don't know.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#906

Earlier quoted context omitted.

Value is entirely in the eye of the beholder - money itself only has value because enough of us agree that it does.

No, it has value because it is the only means that are absolutely accepted to pay taxes and most legal judgements (i.e. obligations to the state that issued that currency.) If you don't have dollars and you need to pay US taxes, you have got to get some or you will go to jail. It's not bitcoin.

You just said the same thing that I did - "enough of us" includes your government.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#908

Earlier quoted context omitted.

The medium obviously matters. One was designed to be cheap and mass produced. Almost anything published has copies in the hundreds or thousands at least and the chances you are concerning yourself with the 'last copy' of some valuable information you can't get elsewhere is minuscle. Film most certainly was not like that.

Assuming these texts exist in large quantities seems to definitionally contradict the term *rare* book. Plenty of books have small print runs or have reached out of print status. One could obviously reproduce films too. Ultimately, you don't know how big N is, and that's what matters. Part of the problem is there is no transparency around this. The population of all books != the population of a specific book or set o…

>Assuming these texts exist in large quantities seems to definitionally contradict the term rare book. Plenty of books have small print runs or have reached out of print status.

They're books that are no longer in print. Doesn't mean there aren't a lot of copies around.

>Ultimately, you don't know how big N is, and that's what matters. Part of the problem is there is no transparency around this. The population of all books != the population of a specific book or set of books. Assuming N is large when you have basically zero information on specifics is foolish. These books might be just as rare as a film for which only one source exists you don't know.

All of this is frankly irrelevant. Books that you buy in bulk at barging bin prices are books that have essentially no market value and were going to the pulp or landfill anyway. Nobody has an obligation to spend money to preserve every extant copy of every book forever. Millions of books are destroyed everyday. Even libraries and archives make these choices. I was simply pointing out the film analogy worked better, but it doesn't change anything.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#909

I don't see any mention of Project Ocean - AKA Google books. Before AI they endevoured to digitize books in a massive online library. This inccluded rare and out of print books many which are archived at libraries. Because they had to preserve the books and return them in the condition they received then they created elaborate technology to accomplish this. The project was met with significant legal challenges from a…

Sounds like a great way for the books to be “preserved” in the hands of someone that will never make them accessible, no thank you Google!

Re: AI companies destroy physical books – let's scan rare books before it's too late

#910

Earlier quoted context omitted.

Whenever I need something from Google Books I inevitably reach the message that this is a limited preview and the part I need is not included. I therefore feel the same way about Google Books that how I felt when I learned that What.cd went down: that I don't gain or lose anything anyway because I never had access to begin with, and that by not making it 100% publicly accessible you're asking for the data to one day…

> Whenever I need something from Google Books I inevitably reach the message that this is a limited preview and the part I need is not included. > I therefore feel the same way about Google Books that how I felt when I learned that What.cd went down: that I don't gain or lose anything anyway because I never had access to begin with, and that by not making it 100% publicly accessible you're asking for the data to one…

Yes, please remember to say “Thank you, Google.”
Post reply on HN