Live data from Hacker News

Archivists Are Trying to Make Sure LibGen Never Goes Down

vice.com

261–270 of 270 posts

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#261
Imagine this:

- A tiny well behaved client that starts with the OS.

- It downloads rare bits of the archive at 1 kb/s obtaining 1 GB every 278 hours. It should stop around 100 MB to 5 GB.

- It periodically announces what chunks/documents it has.

- It seeds those chunks at 1 kb/s

- Chunks/documents that have thousands of seeds already are not announced. Eventually those are pruned.

This escalates the situation to the point where everyone can help without it costing anything.

If someone is trying to obtain a 20 mb pfd it would take 5 and a half hours using a single 1 kb seed. With just 50 seeds it's just 8 min.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#262
post #257

Earlier quoted context omitted.

Because those hundreds of years don't transpire in a glimpse. At some point in the middle there will be deprecated formats and new ones, and transcoders you can batch run. Sure it relies on intervention, but the upside is any/everyone else can copy the one persons work. Yes we should learn from history, but we should also not assume that everything that happened before will happen the same way again, given how much o…

> However, without archivists actively transforming content to new formats as required, it might only take a few decades before a lot of content starts to require a massive effort to read.

More effort than batch reading physical books and tablets in old languages?

You can reuse interfaces easier on data, and current ML could probably pull some of the weight of interpreting old data right now, not to mention what we have 50 years from now.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#263

Earlier quoted context omitted.

No one is proposing we use floppy disks. Redundant, shared servers ARE a forever solution. Making sure your data is one one of the ones that makes it seems like a vastly easier proposition to me than writing data to clay tablets and trying to keep those from ending up in a dump somewhere.

What is the likelihood that historians a century or two hence will have an application capable of turning an ISO 32000-1 file into a human-readable text? If we are talking about archaeologists, rather than historians, even ASCII and Unicode could be a challenge to work out.

0.99999 at least.

Compare the capabilities of digital historians today to those 10- and 20-years ago respectively. It’s night and day.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#264

Earlier quoted context omitted.

Can you elaborate? What's the catch?

The only catch is that it's a minor lottery which model drive you're getting. For instance, I got all white label WD80EMAZs (256MB cache, non-SMR, same firmware as the Reds) in this batch, so I had to insulate the 3.3V pins. There are also true Reds, 128MB and 512MB cache drives, helium filleds, 7.2K HGSTs slowed to 5.4K, and other variants.

Or use a traditional power supply to sata cable.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#265

Earlier quoted context omitted.

This is a very restricted subset of utf-8. I agree that the ASCII subset would not be tremendously difficult to decipher; the most interesting parts are laid out systematically and in order and case is even just a bit flip. It's even fairly plausible that the utf-8 numerical encoding can be reverse-engineered from a few samples; enough languages' text generally only use characters from few enough blocks to identify.…

I agree recovering CJK Unified Ideographs encodings would be far harder than a phonetic alphabet, however a few things could make not as hard as it seems. The decoder has access to a text in both the future format and UTF-8. A text might mix phonetic words and ideographs as Japanese sometimes does today. The phonetic words would provide clues as to the ideographic characters. Code breakers have decoded ciphertexts wh…

Exactly. You shouldn't underestimate the tremendous amount of work has been put into deciphering actual ancient languages using advanced techniques and minor contextual clues. Compared to that, deciphering most common UTF8 data would be relatively simple, meaning it could be done by a single person with some reverse engineering skills.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#266
post #2

This is an extremely important effort. The LibGen archive contains around 32 TBs of books (by far the most common being scientific books and textbooks, with a healthy dose of non-STEM). The SciMag archive, backing up Sci-Hub, clocks in at around 67 TBs [0]. This is invaluable data that should not be lost. If you want to contribute, here's a few ways to do so. If you wish to donate bandwidth or storage, I personally k…

Mind explaining the origin of your 32 TB figure? I must be missing something enormous, but as far as I can tell the SciMag database dump is 9.3 GB, the LibGen non-fiction dump is 3.2 GB, and the LibGen fiction dump is 757 MB. That's a pretty huge divergence.

Source: http://gen.lib.rus.ec/dbdumps/

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#267
post #2

This is an extremely important effort. The LibGen archive contains around 32 TBs of books (by far the most common being scientific books and textbooks, with a healthy dose of non-STEM). The SciMag archive, backing up Sci-Hub, clocks in at around 67 TBs [0]. This is invaluable data that should not be lost. If you want to contribute, here's a few ways to do so. If you wish to donate bandwidth or storage, I personally k…

Mind explaining the origin of your 32 TB figure? I must be missing something enormous, but as far as I can tell the SciMag database dump is 9.3 GB, the LibGen non-fiction dump is 3.2 GB, and the LibGen fiction dump is 757 MB. That's a pretty huge divergence. Source: http://gen.lib.rus.ec/dbdumps/

Oh, wait. I'm dumb. I see that your first link is a citation.

Continuing to be dense, why is there a difference between their "database dump" and the total of all the files they have?

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#268
post #2

This is an extremely important effort. The LibGen archive contains around 32 TBs of books (by far the most common being scientific books and textbooks, with a healthy dose of non-STEM). The SciMag archive, backing up Sci-Hub, clocks in at around 67 TBs [0]. This is invaluable data that should not be lost. If you want to contribute, here's a few ways to do so. If you wish to donate bandwidth or storage, I personally k…

Mind explaining the origin of your 32 TB figure? I must be missing something enormous, but as far as I can tell the SciMag database dump is 9.3 GB, the LibGen non-fiction dump is 3.2 GB, and the LibGen fiction dump is 757 MB. That's a pretty huge divergence. Source: http://gen.lib.rus.ec/dbdumps/

The databases contain the metadata (authors, edition, ISBN, etc.) for the books.

Thus, 32 TB of books (over 2 million titles), 3.2 GB database.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#269
post #258

Earlier quoted context omitted.

Nature shows us how to process information at ever increasing noise and scale - https://www.edge.org/response-detail/10464

Yes and no. Briefly: the article distinguishes "endocrinal" vs. "distributed" decisionmaking. This applies at some levels, but not at others. For individual humans, we don't have the option of rewiring our concsiousnesses, which are rather pathetically single-threaded, and can at best multitask poorly by task-switching, at a very great loss of task proficiency. Even withing collective organisations (companies, govern…

The article "Evolving the Global Brain" was thought-provoking, especially in the context of our discussion about the history of information and the exponentially increasing amount of information for humanity to gather/produce, process, curate, archive.

It's an attractive concept, that human society is structurally similar to a brain, and that an individual is a neuron. (If humanity is the brain, I suppose the rest of the Earth is the body. We're not doing too well as the self-appointed brain of the operation.)

My first reaction to the analogy of "endocrinal" (one-to-many) and "neural" (many-to-many) decision making, is that it's missing a primal psychological/biological motivation of humans to seek to dominate others of its own kind as well as all of nature. I'm not familiar enough with biology to say definitively, but I'm pretty sure the endocrinal system does not actively seek to subjugate the neural system (or vice versa) and dominate the whole body.

Social organization, it seems to me, is more a function of power, very small groups gaining advantage and dominance over vastly larger groups of people, than that of collaboration for mutual benefit. (I might be a bit too cynical of political motivations and authentic democracy these days.)

From the final paragraph:

> ..the current global brain is only tenuously linked to the organs of international power. Political, economic and military power remains insulated from the global brain, and powerful individuals can be expected to cling tightly to the endocrine model of control and information exchange.

I'd disagree with this, and say that the global brain (if we mean the Internet and its empowerment of globally networked intelligence) was born from the wombs of "political, economic and military power". It never achieved escape velocity to become a truly free, autonomous and collaborative, neural model of decision making.

To backtrack a bit:

> Well-connected collective entities like Google and Wikipedia will play the role of brainstem nuclei to which all other information nexuses must adapt.

The most powerfully well-connected collective entities are international political/financial/corporate entities, and indeed do they more or less dictate how all information nexuses (nexii?) must adapt.

One biological analogy that comes to mind, is how propaganda and "disinformation" act like neurotoxins in the social brain, introducing noise/entropy, skewing its coherence, and preventing well-informed and orchestrated cooperation.

Another is how established political powers have a well-developed "immune system", composed of mass media, legal structures, military/police force, surveillance of the public. This immune system could be seen at work, for example, at the environmental protests at the Standing Rock Indian Reservation.

The final sentence of the article:

> This formidable design task is left up to us.

By this I assume the author means, evolving the global brain. Quite a challenge! From my perspective, it's going to be a historic struggle: design or be designed.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#270
post #268

Earlier quoted context omitted.

Mind explaining the origin of your 32 TB figure? I must be missing something enormous, but as far as I can tell the SciMag database dump is 9.3 GB, the LibGen non-fiction dump is 3.2 GB, and the LibGen fiction dump is 757 MB. That's a pretty huge divergence. Source: http://gen.lib.rus.ec/dbdumps/

The databases contain the metadata (authors, edition, ISBN, etc.) for the books. Thus, 32 TB of books (over 2 million titles), 3.2 GB database.

Ah, that makes sense.

To make sure I'm understanding this correctly:

The Libgen Desktop application (which requires only a copy of the database) would then use the DB metadata to make LibGen locally searchable, and would only retrieve the individual books/papers on request?

Post reply on HN