there are non destructive scanning options but they are nowhere as fast and prone to errors
AI companies destroy physical books – let's scan rare books before it's too late
721–730 of 961 posts
Re: AI companies destroy physical books – let's scan rare books before it's too late
#722I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collective responsibility. Thus, the preservation of this knowledge, its availability, and accessibility are t…
>From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge. How tho? Ever sunday at the flea market, I see thousands of books that are rotting, hoping for someone to buy them or at least take t…
Re: AI companies destroy physical books – let's scan rare books before it's too late
#723I wholeheartedly believe the AI controversy on destroying books is being stirred up by the companies themselves. Copyright law requires you destroy a book, if you format shift it. If you digitise, you need to ensure its not a "copy" but that your one license went with the book. So... If enough people complain, they get to pressure for copyright changes. Which will just so happen to have massive carveouts to let them…
"You own a particular physical copy; you don't possess an abstract transferable 'one-copy license." There is nothing that says you have to destroy something because you scanned it. This argument has been confusing me since I've seen this pop up. Edit; despite the above , looking at the court documents from the Anthropic case, this is pretty close to what they were arguing: “we are just transferring the physical form…
It's also probably more convenient for them to destroy the books as opposed to trying to find space to store them. Knowing those companies, most likely they'd be just stuffed into some warehouse to rot after a couple years.
Re: AI companies destroy physical books – let's scan rare books before it's too late
#724Earlier quoted context omitted.
> a digital copy with the ability of doing millions of copies is stored somewhere somewhere were we can't access it. as the article states: “permanently locking human knowledge inside private corporate servers” the issue isn't that the physical copy is gone, it's that they are preventing people from making digital copies that are actually accessible by destroying the physical copies. > If you go to a recycling centre…
> somewhere were we can't access it By any metric imaginable, it's making the information more accessible, not less. First, it's taking a single copy of a 10k physical print and it's making it digital. Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed. But it's _closer_ to being widely available, not farther. Plus having the info part of a LLM makes it immediately available to…
I don't if that is true: a lot of old books might still have copy or other rights associated to them, likely owned by author and/or publisher, directly or inherited, but often those who have the rights do not have digital or physical copies at hand anymore (some old books are, well, really old). Does Anthropic make sure to track down, contact and then share the digital copy they make with those who have rights on the work? If not, they are not making it in any way easier to re-print the books, while making their supply more scarce (they destroy existing embodiments).
Re: AI companies destroy physical books – let's scan rare books before it's too late
#725Re: AI companies destroy physical books – let's scan rare books before it's too late
#726Earlier quoted context omitted.
Nuclear killed itself (vast cost overruns); there was no need for hallucinated foreign influences. But I understand blaming your energy waifu for its own failure is unacceptable for nuclear bros.
nuclear was killed by russian nat gas money. running nuclear plants were shut down while running just fine
What other nonsense do you believe?
Re: AI companies destroy physical books – let's scan rare books before it's too late
#727It is not a big deal. Since the invention of the printing press any important book has been duplicated by thousands, tens of thousands or even million of units. Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is stored somewhere) is not problematic for Humanity. By the way, I always search for second hand books. Most of the books there ar…
It is common for academic books to have publication runs in the low three digits.
You may argue these books are not important. But how do we know if we fail to preserve it?
Re: AI companies destroy physical books – let's scan rare books before it's too late
#728Re: AI companies destroy physical books – let's scan rare books before it's too late
#729I don't see any mention of Project Ocean - AKA Google books. Before AI they endevoured to digitize books in a massive online library. This inccluded rare and out of print books many which are archived at libraries. Because they had to preserve the books and return them in the condition they received then they created elaborate technology to accomplish this. The project was met with significant legal challenges from a…
Like all things, Google will eventually realize they cannot make significant ad revenue and they will eventually give up and discontinue serving this, though I doubt it's more than a scratch in terms of disk space. It's great they did this, but the Google that is today cannot be trusted with data of public value anymore.
Re: AI companies destroy physical books – let's scan rare books before it's too late
#730Earlier quoted context omitted.
>Plus having the info part of a LLM makes it immediately available to literally billions. Help me understand how. Not only are these LLMs expressly prohibited from specifically regurgitating copyright works if the users asks them to, but they habitually hallucinate or paraphrase things wrong. If they won't regurgitate the copyrighted text verbatim, and are known to be confidently incorrect and hallucinatory, I'm stru…
Depends upon what you want. For example, knowledge about how the book smells when you open it, something that reader do talk about enough I don't think this should be a strawman, is lost. But, that is about the experience of reading the book, not the knowledge of the book. The exact text? Yeah, I think that is largely lost as well. This is a summary. And for rarer books, it will be a particularly bad summary. The bas…
No, it’s not. Undigitized data is still data. Wording, style, nuance, type, artwork, binding, metadata, attributions, citability… all of these things are permanently lost, which is a fucking tragedy because they don’t have to be. Even if you’re not willing to take the time to scan every page in a cradle scanner, as rare books should be, (and don’t tell me they don’t have the money to get a handful of library interns to do this,) you can disbind the books and store them as they did in the Caselaw Access Project at Harvard Law. They removed the pages from the binding, scanned them on a high speed conveyor belt scanner which yielded full color 600 DPI jp2 images, placed the pages back in the binding like a folio that could be re-bound if needed, vacuum sealed them, and stored them in a salt mine. It’s not like it was slow, either — we did 40k in 18 months and we did take the time to scan the rare ones with a cradle scanner. And we did it all in less than open AI probably spends in a day on inference.
> So, in that sense, the knowledge is better being spread compared to copyright where the book stays in a warehouse until it is disposed of.
That’s a false dichotomy. Libraries exist for this exact reason, and their not already having a copy does not make “you snooze you lose” a morally acceptable strategy.
I’ve been pretty cool on the direction of SV for the past decade at least, but I am absolutely gobsmacked by the unbridled hubris of these companies over the past 5 years.