Earlier quoted context omitted.
It’s incredibly sad: https://www.theatlantic.com/technology/archive/2017/04/the-t...
There are also multiple petabytes of microfiche scans of old newspapers. And of course nobody cares about it. The problem was shut down 2011ish and the data became "owned" by a team that didn't care for it. There was talk of just deleting the data because the team didn't want to pay for it. Ugh.
Archivists Are Trying to Make Sure LibGen Never Goes Down
171–180 of 270 posts
Re: Archivists Are Trying to Make Sure LibGen Never Goes Down
#172What's interesting is that 32TB is becoming more and more affordable and the research material is roughly staying about the same size. That might change though as people start including video + data within papers and have new notebook formats that are live and contain docker containers/ipython, etc. It's a shame we can't just mail these around.
Re: Archivists Are Trying to Make Sure LibGen Never Goes Down
#173Earlier quoted context omitted.
I worry that if this system becames permanent, one in which it is practically impossible to stop piracy, followed by the loss of traditional incentives we might find ourselves in a place where no motivated investor will break even when producing quality and innocuous content.
In that case, we will just pass around the same old stuff until we get bored enough that we'll actually pay for new stuff.
Re: Archivists Are Trying to Make Sure LibGen Never Goes Down
#174The new architecture of pirate sites, what I call the Hydra architecture, seems pretty interesting to me. There isn't a single site hosting the content, but a group of mirrors freely exchanging data between one another. In case some of them go down, the other ones still remain and new ones can appear, copying data from the remaining mirrors. This is like a hydra that grows two heads every time you chop one off. It's…
> It's absolutely unkillable Just like any other distributed system, this is vulnerable to organized take downs and scare tactics. There was a whole bunch of mirrors of Pirate Bay, yet once most of Europe's legal systems adopted the "sharing is theft" mindset, it became pretty much impossible to find one.
Re: Archivists Are Trying to Make Sure LibGen Never Goes Down
#175Earlier quoted context omitted.
I read the Wikipedia article about it and the sad thing is that the majority of the Yongle Encyclopedia seem to have been destroyed only in quite recent times.
> but 90 percent of the 1567 manuscript survived until the Second Opium War in the Qing dynasty. In 1860, the Anglo-French invasion of Beijing resulted in extensive burning and looting of the city,[16] with the British and French soldiers taking large portions of the manuscript as souvenirs. Preservation is easy if you don't get invaded.
Re: Archivists Are Trying to Make Sure LibGen Never Goes Down
#176Libgen is one of the greatest contributors to scientific productivity worldwide, possibly beaten only by Sci-Hub. Just about everybody in academia knows about it. If it ever vanished, some of us could probably still get by trading files from person to person, but nothing could be as perfect as what we got now.
Re: Archivists Are Trying to Make Sure LibGen Never Goes Down
#177Earlier quoted context omitted.
I think the duplication issue is probably overstated. I doubt tackling that would shave off more than 20% of the total backup size.
Speaking from personal experience, I usually see several results for any search. Granted, there's a big selection bias there, but 20% seems way too small.
Re: Archivists Are Trying to Make Sure LibGen Never Goes Down
#178Earlier quoted context omitted.
> It's absolutely unkillable Just like any other distributed system, this is vulnerable to organized take downs and scare tactics. There was a whole bunch of mirrors of Pirate Bay, yet once most of Europe's legal systems adopted the "sharing is theft" mindset, it became pretty much impossible to find one.
But now the main site seems to be bullet proof. There was a time where weekly there would be a new official link. I'm not sure what changed structurally with hosting tbp
Re: Archivists Are Trying to Make Sure LibGen Never Goes Down
#179Maybe we should print this out on acid-free paper-thin flexible wood-pulp sheets stitched to together to form linear organized aggregations. Each aggregation would contain one or more works and be searchable using a SQL-like database. To make this plan really work there would need to be a collection of geographically distributed long term physical repositories that would receive periodic updates as new material becam…
If you found a mysterious archive object and had no idea what it was - CD-R, hard drive, SSD, whatever - not only would you have to reinvent an entire hardware reader around it, you would also have to work out the file structure, extract the data (some of which could be damaged), and reverse engineer the container file formats and the data structures inside them.
If you got all of that right, you'd eventually be able to start trying to translate the content of the text, audio, images, videos (how many compression formats are there?) into something you could understand.
A much more advanced civilisation would struggle with making a cold start on all of that. In our current state, we'd get nowhere if we didn't already have some records explaining where to begin.
Re: Archivists Are Trying to Make Sure LibGen Never Goes Down
#180Earlier quoted context omitted.
> I do wonder wither digital or analogue formats are better able to survive into the distant future. There are 5000 year old clay tablets we can still read. There are centuries old documents on paper, vellum etc. that we can still read. I personally have decades-old paper documents I can easily read, and a box of floppies I can't. It's not just a problem of unreadable physical media, I have a database file on a perfe…
Clay is the plastic of the ancient world. Let's say the probability that: a single copy of a physical book survives 1,000 years, is found and is understood by an archaeologist , is pB and the probability that a single copy of a book on an SSD survives 1,000 years is found and understood by an archaeologist is pD. Even if pB is far larger than pD it could be the case that there might be so many more copies of single b…
Yes. That's what I mean by LOCKSS being easier.
> is found and is understood by an archaeologist,
There is a problem with merging these two probabilities.
The probability of finding a book is of course massively smaller than the probability of finding a digital copy.
The probability of understanding a book is so much greater than the probability of understanding a file on a disk.
This makes it more likely that the physical book will survive in a meaningful way.
> It could also be the case that each generation would copy these books onto new digital media
This is what I mean by archivists actively transforming the content. Regarding written content like the Iliad, copies and translations can be made centuries apart. Content in digital formats may need to be transformed whenever the application that reads it is discontinued.