Live data from Hacker News

Archivists Are Trying to Make Sure LibGen Never Goes Down

vice.com

171–180 of 270 posts

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#171
post #40

Earlier quoted context omitted.

It’s incredibly sad: https://www.theatlantic.com/technology/archive/2017/04/the-t...

There are also multiple petabytes of microfiche scans of old newspapers. And of course nobody cares about it. The problem was shut down 2011ish and the data became "owned" by a team that didn't care for it. There was talk of just deleting the data because the team didn't want to pay for it. Ugh.

Huh interesting. Have you got any details? Names of the project? There are probably institutions willing to host that content.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#172

What's interesting is that 32TB is becoming more and more affordable and the research material is roughly staying about the same size. That might change though as people start including video + data within papers and have new notebook formats that are live and contain docker containers/ipython, etc. It's a shame we can't just mail these around.

When people publish data it's typically uploaded to a public repository anyway. Supplementary videos are a thing, but in my field at least they generally stay in the supplementary and aren't the raw data so file sizes are reasonable, while still images are used in the text. Journals are still printed works first, believe it or not.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#173
post #159

Earlier quoted context omitted.

I worry that if this system becames permanent, one in which it is practically impossible to stop piracy, followed by the loss of traditional incentives we might find ourselves in a place where no motivated investor will break even when producing quality and innocuous content.

In that case, we will just pass around the same old stuff until we get bored enough that we'll actually pay for new stuff.

Some people might decide to pay but the technology will be there to distribute it for free. At this point it would be sort of like a public good with the free rider problem.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#174

The new architecture of pirate sites, what I call the Hydra architecture, seems pretty interesting to me. There isn't a single site hosting the content, but a group of mirrors freely exchanging data between one another. In case some of them go down, the other ones still remain and new ones can appear, copying data from the remaining mirrors. This is like a hydra that grows two heads every time you chop one off. It's…

> It's absolutely unkillable Just like any other distributed system, this is vulnerable to organized take downs and scare tactics. There was a whole bunch of mirrors of Pirate Bay, yet once most of Europe's legal systems adopted the "sharing is theft" mindset, it became pretty much impossible to find one.

But now the main site seems to be bullet proof. There was a time where weekly there would be a new official link. I'm not sure what changed structurally with hosting tbp

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#175
post #113

Earlier quoted context omitted.

I read the Wikipedia article about it and the sad thing is that the majority of the Yongle Encyclopedia seem to have been destroyed only in quite recent times.

> but 90 percent of the 1567 manuscript survived until the Second Opium War in the Qing dynasty. In 1860, the Anglo-French invasion of Beijing resulted in extensive burning and looting of the city,[16] with the British and French soldiers taking large portions of the manuscript as souvenirs. Preservation is easy if you don't get invaded.

It's easy if you anticipate these things. Who put the dead sea scrolls in that cave in the middle of nowhere? Not someone who went in and forgot their scroll one day. Someone who had the foresight that this would be a safe space in the face of who knows what future threat. And it payed off.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#176
post #130

Libgen is one of the greatest contributors to scientific productivity worldwide, possibly beaten only by Sci-Hub. Just about everybody in academia knows about it. If it ever vanished, some of us could probably still get by trading files from person to person, but nothing could be as perfect as what we got now.

It's saved me probably 3 grand over the course of college. I would have had to take on debt otherwise to pass my courses.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#177

Earlier quoted context omitted.

I think the duplication issue is probably overstated. I doubt tackling that would shave off more than 20% of the total backup size.

Speaking from personal experience, I usually see several results for any search. Granted, there's a big selection bias there, but 20% seems way too small.

In my experience it's different editions or mirrors.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#178
post #174

Earlier quoted context omitted.

> It's absolutely unkillable Just like any other distributed system, this is vulnerable to organized take downs and scare tactics. There was a whole bunch of mirrors of Pirate Bay, yet once most of Europe's legal systems adopted the "sharing is theft" mindset, it became pretty much impossible to find one.

But now the main site seems to be bullet proof. There was a time where weekly there would be a new official link. I'm not sure what changed structurally with hosting tbp

They just stopped going after it, and focused resources on stopping streaming websites

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#179

Maybe we should print this out on acid-free paper-thin flexible wood-pulp sheets stitched to together to form linear organized aggregations. Each aggregation would contain one or more works and be searchable using a SQL-like database. To make this plan really work there would need to be a collection of geographically distributed long term physical repositories that would receive periodic updates as new material becam…

This is not a solvable problem without technological continuity, or some unimaginably smart technology we can't imagine today.

If you found a mysterious archive object and had no idea what it was - CD-R, hard drive, SSD, whatever - not only would you have to reinvent an entire hardware reader around it, you would also have to work out the file structure, extract the data (some of which could be damaged), and reverse engineer the container file formats and the data structures inside them.

If you got all of that right, you'd eventually be able to start trying to translate the content of the text, audio, images, videos (how many compression formats are there?) into something you could understand.

A much more advanced civilisation would struggle with making a cold start on all of that. In our current state, we'd get nowhere if we didn't already have some records explaining where to begin.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#180

Earlier quoted context omitted.

> I do wonder wither digital or analogue formats are better able to survive into the distant future. There are 5000 year old clay tablets we can still read. There are centuries old documents on paper, vellum etc. that we can still read. I personally have decades-old paper documents I can easily read, and a box of floppies I can't. It's not just a problem of unreadable physical media, I have a database file on a perfe…

Clay is the plastic of the ancient world. Let's say the probability that: a single copy of a physical book survives 1,000 years, is found and is understood by an archaeologist , is pB and the probability that a single copy of a book on an SSD survives 1,000 years is found and understood by an archaeologist is pD. Even if pB is far larger than pD it could be the case that there might be so many more copies of single b…

> thus making it more likely the book will survive via an SSD than a physical book

Yes. That's what I mean by LOCKSS being easier.

> is found and is understood by an archaeologist,

There is a problem with merging these two probabilities.

The probability of finding a book is of course massively smaller than the probability of finding a digital copy.

The probability of understanding a book is so much greater than the probability of understanding a file on a disk.

This makes it more likely that the physical book will survive in a meaningful way.

> It could also be the case that each generation would copy these books onto new digital media

This is what I mean by archivists actively transforming the content. Regarding written content like the Iliad, copies and translations can be made centuries apart. Content in digital formats may need to be transformed whenever the application that reads it is discontinued.

Post reply on HN