Live data from Hacker News

AI companies destroy physical books – let's scan rare books before it's too late

annas-archive.gl

211–220 of 961 posts

Re: AI companies destroy physical books – let's scan rare books before it's too late

#211

Earlier quoted context omitted.

But that scan is never made available to us in its original form. So it getting scanned by the AI company does nothing to preserve the book.

Dumpsters also don't typically come equipped with a robot scanner and network uplink built in. Like, I really don't know what people objecting to this imagine typically happens to old, unwanted books. They don't get sent to some magical library in the countryside if unpurchased where they are carefully maintained forever (next to where Rover spends the rest of his days). They are very literally thrown into the trash.…

The Internet Archive tries to be that magical library, but they can only scan and physically archive what is sent to them.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#212

Earlier quoted context omitted.

Can you get the exact text back out with a prompt or not? Having or not having a book isn't fuzzy.

Having or not having a book is absolutely fuzzy. If you have a translation, do you have the book? Even if, like the Odyssey, there are hundreds of wildly varying translations? What about an abridged copy? What about the Sparknotes version? If you have a copy of Pride And Prejudice And Zombies, do you have a copy of Pride And Prejudice? Certainly more so than if you have neither.

>If you have a translation, do you have the book?

No, you have a translation.

>Even if, like the Odyssey, there are hundreds of wildly varying translations?

Precisely why translations are not considered equivalent to the original text.

>What about an abridged copy? What about the Sparknotes version?

An abridged copy is not a copy of the unabridged version.

>If you have a copy of Pride And Prejudice And Zombies, do you have a copy of Pride And Prejudice?

No.

I'm honestly surprised these were the questions you chose to ask, when you could have asked what if you have 90% of the pages, or what if most of the pages are missing pieces because the book was shot with a shotgun, or what if the book was scanned and OCRed and all the "rn"s were replaced with "m"s and all the lower case Ls with ones. Hell, is a scan of the book close enough to having the book, or is it far enough that one can no longer be said to have the book anymore?

Re: AI companies destroy physical books – let's scan rare books before it's too late

#213
post #87

As much as I hate piracy in a sector in financial crisis like book publishing (because Anna’s project is piracy), I hate even more what these large AI companies are doing: privatizing human knowledge. On one side, there’s copyright law, which exists to support the work of creative people. “Information wants to be free” is bullshit spread by people who have never spent a minute in their lives trying to create somethin…

> physical book sales are plummeting — the main source of income for writers — and shadow libraries are killing the incentive to write

there have been many times I've wanted to pay for an ebook but found that the only places that sell it apply DRM to it, so shadow library it is

Re: AI companies destroy physical books – let's scan rare books before it's too late

#214
post #189

Earlier quoted context omitted.

They are, they are in no risk of being arrested. But that does not make it free of moral judgement.

Why not?

That's just a fact, not something to argue around. People will judge you for many reasons, some cultural, some political, some ethical, some personal, some valid, some not. Its human nature.

Which is also not illegal and within some bounds and exceptions, a protected right across the globe.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#215
post #155

Earlier quoted context omitted.

And they should not even be needed. In many places the issue is solved at start. Copy or copies of each commercially produced book is send to national library. Which with tax payer money keeps an archive. Meaning that at least one copy exist for research purposes if needed.

Unfortunately, that's not the case in the United States. The LOC only selects around half of published books to be permanently held. The rest are disposed of (usually returning them to the publisher, donating them to a library, or destroying them).

They should send them to The Internet Archive instead.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#216

The piracy organizations are playing 4D chess while everyone else is playing checkers. The irony of this entire situation - AI companies being legally required to shred books due to kafkaesque copyright laws, then used as a marketing tactic by Anna's Archive - is a work of art. I support Anna's Archive, by the way. Information wants to be free.

Plenty of written works aren’t “information” but rather art. Most piracy is just about people preferring not to pay for novels, TV and film.

There is also a ton of tv shows, movies, music and books that you cannot buy, for now real good reason. I wouldn't be surprised that if in a few years there will be shows and movies that are only exists as pirated versions. With things increasingly only being available on streaming platforms or behind DRM in other ways, we risk looking back on the current era as a black hole 50 years from now.

My concern is that less popular content is just erases, lost in mergers or lost in massive datacenters, never to be seen again.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#217

Earlier quoted context omitted.

...because it is illegal to copy copyrighted material. 70 years later they might do it.

The thing I find most hypocritical though is that they are probably never share their libraries with anyone. After scanning, downloading, stealing, overloading websites and everything in between, to acquire enough data for their stupid machine, they're not going to share their data? I get that most of it can't be shared, but a lot can. There's no reason why you need to destroy multiple copies of a book from 1880, whe…

Since they seem to leapfrog each others’ models every few months, the training data is one of the few ways they can build competitive advantage, and that explains why they don’t share, even if we don’t have to like this.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#218

Earlier quoted context omitted.

...because it is illegal to copy copyrighted material. 70 years later they might do it.

The thing I find most hypocritical though is that they are probably never share their libraries with anyone. After scanning, downloading, stealing, overloading websites and everything in between, to acquire enough data for their stupid machine, they're not going to share their data? I get that most of it can't be shared, but a lot can. There's no reason why you need to destroy multiple copies of a book from 1880, whe…

This is solved by law, which is solved by 'we the people' and I bet many AI companies would be fine with something like the equivalent to patent law with bankruptcy escrow to the library of congress, where they must release the scans in 10 years for books that the vast majority will not give a flying shit about. By then the advantage is long gone in data moat.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#219
post #87

As much as I hate piracy in a sector in financial crisis like book publishing (because Anna’s project is piracy), I hate even more what these large AI companies are doing: privatizing human knowledge. On one side, there’s copyright law, which exists to support the work of creative people. “Information wants to be free” is bullshit spread by people who have never spent a minute in their lives trying to create somethin…

>In a world like that, what motivation would we still have to read? Why would a reduction in human writers cause a complete reduction in motivation to read? There's millions of books already written and it makes zero sense that people would stop writing. People write for hundreds of reasons other than to make money and they created literature before copyright was a thing.

First, people want tonread books that speak to them and their lives, older books are not that.

Second, people write to be read. It takes huge amount of effort too. With no potential reward for it at all, they stop.

Third, we are social animals. If you dont see people reading, if you dont read yourself, you wont even think of writing.

Re: AI companies destroy physical books – let's scan rare books before it's too late

#220
post #88

Earlier quoted context omitted.

Your what now? Made from your personal waste? That does sound unique. https://en.wiktionary.org/wiki/wastebook Oh right. But anyway, nobody knows what needs preservation, it's a basic problem of life, somebody usually mentions the BBC throwing out boring old Doctor Who tapes to save archive space because nobody liked it any more at that point in time. Some things should probably be thrown out now and then, I suppose.

The alternative for many of these un(der)appreciated books is that they will get unceremoniously dumped in the future anyway. The publishing industry and libraries etc dispose off lots and lots of books. So at least with the AI companies they are scanning them and preserving them digitally. Not just in the trained weights, but also as raw training data for future runs. P.S. I'm not sure why you need to make fun of yo…

The AI companies’ working assumption is that if someone found it worth printing, it has enough information content to help train a model. That assumption might be invalid with some of the more rambling self-published books, however.
Post reply on HN