Live data from Hacker News

Archivists Are Trying to Make Sure LibGen Never Goes Down

vice.com

251–260 of 270 posts

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#251

Earlier quoted context omitted.

Bitmaps are easy enough, but I wouldn't bet on UTF-8. And any modern compression is probably right out without technological continuity.

I think if you gave a philologist living in 1880 AD a clay tablet with a binary inscription of a fragment of an English poem encoded UTF-8 they would decode it very quickly. This is what the philologist would see: >...ABABABBBABBABAAAABBBBAABAABABBAAAABAAAAAABBABAABABBAABBAAABAAAAAAABAABBBABBBABAAABBABAABABBBAABBAABAAAAAABBAABAAABBAAAABABBABBBAABBAAABBABBABAABABBABBBAABBAABBBAABAAAAAABBBBAABABBABBBBABBBABABAABAAAAAAB…

This is a very restricted subset of utf-8. I agree that the ASCII subset would not be tremendously difficult to decipher; the most interesting parts are laid out systematically and in order and case is even just a bit flip.

It's even fairly plausible that the utf-8 numerical encoding can be reverse-engineered from a few samples; enough languages' text generally only use characters from few enough blocks to identify. If you're really motivated, you can probably work your way through most of the languages with phonetic writing systems.

But then there's CJK Unified Ideographs, where the characters that get used are scattered essentially randomly because the ordering is only relevant if you already know how many and which characters were encoded at what point in the history of Unicode.

There are large swaths of Unicode which, if somehow totally lost, would essentially require finding font data or character reference tables to recover.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#252

Earlier quoted context omitted.

This isn't even movies wherein some large studio's can send notices. I don't think publishing houses have that many funds to send so many legal notices. Books are a safe bet to pirate

That's what people said about music and films too. You don't want to be the next Jammie Thomas. This is an existential threat to the deep-pocketed likes of Elsevier et al. They will use the law to make an example of anyone too close to their sphere of influence; so if you are in the US or the EU; support the efforts of LibGen vocally and loudly, and contribute anonymously, but don't risk your neck to the extent where…

I think a turn-key solution for people living in not US/EU will still help the general health of the archive.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#253
post #231

Earlier quoted context omitted.

What does that have to do with sharing knowledge in the USSR and the countries in the Soviet block? It was never in Soviet ideology to hide knowledge behind paywalls. See, for example, this [0] post about Mir publishing house and warm comments of Indians who grew up with their books. Sci-hub's ideology is just continuation of this approach. [0] https://news.ycombinator.com/item?id=21352277

> Sci-hub's ideology is just continuation of this approach. Actually, that was the point of what I was saying—the mathematicians had to be inventive and thus passed around preprints that they knew would also be read in the West.

Why did they have to be inventive? Please provide a source.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#254

Earlier quoted context omitted.

I think if you gave a philologist living in 1880 AD a clay tablet with a binary inscription of a fragment of an English poem encoded UTF-8 they would decode it very quickly. This is what the philologist would see: >...ABABABBBABBABAAAABBBBAABAABABBAAAABAAAAAABBABAABABBAABBAAABAAAAAAABAABBBABBBABAAABBABAABABBBAABBAABAAAAAABBAABAAABBAAAABABBABBBAABBAAABBABBABAABABBABBBAABBAABBBAABAAAAAABBBBAABABBABBBBABBBABABAABAAAAAAB…

This is a very restricted subset of utf-8. I agree that the ASCII subset would not be tremendously difficult to decipher; the most interesting parts are laid out systematically and in order and case is even just a bit flip. It's even fairly plausible that the utf-8 numerical encoding can be reverse-engineered from a few samples; enough languages' text generally only use characters from few enough blocks to identify.…

I agree recovering CJK Unified Ideographs encodings would be far harder than a phonetic alphabet, however a few things could make not as hard as it seems. The decoder has access to a text in both the future format and UTF-8. A text might mix phonetic words and ideographs as Japanese sometimes does today. The phonetic words would provide clues as to the ideographic characters.

Code breakers have decoded ciphertexts which used a code such that each word was replaced with a number. To make it even harder common words would be replaced by more than one numbers to defeat common frequency analysis techniques. This was done often with pen and paper.

Yuri Knorozov managed to decipher the Mayan script. That was a significantly harder task than recovering UTF-8 mappings because he has very little to work with on the source language (he did have somethings).

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#255

Earlier quoted context omitted.

It is very common these days to buy the WD 8TB, 10TB and 12TB external USB3 hard drives and remove their cases, and put them in some sort of home built file server or NAS. There's a technique to put a thin section of kapton tape on one of the SATA pins so that they will power up from ordinary PC/ATX type power supplies with regular SATA power connectors. https://www.instructables.com/id/How-to-Fix-the-33V-Pin-Issu...…

> at no greater or lesser annual failure rate than the expensive enterprise hard drives. I've read these reports as well, but I can say that it's not my experience (we've gone through a few rounds of shucking at the Internet Archive, for economy and in one case necessity after the 2011 Thailand floods pinched the supply chain). Our raw failure rates on shucked drives are significantly higher, and the drives themselve…

Could you publish some statistics?

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#256

Earlier quoted context omitted.

Bush's essay is of course a classic. There are some precursors -- there's a BBC interview of H.G. Wells describing something similar from the 1940s.[1] E.F. Forster's The Machine Stops has some similar ideas. And various encyclopaedists very much embodied similar ideals. I've been listening to Peter Adamson's "History of Philsophy Without Any Gaps" podcast, which is excellent , and spends a fair bit of time looking a…

While brushing up on the encyclopaedists, I found this little gem: "Among some excellent men, there were some weak, average, and absolutely bad ones. From this mixture in the publication, we find the draft of a schoolboy next to a masterpiece." — Denis Diderot Taking the quote out of context (and aside from its historical male-centered language) - it sure rings true of the current state of the web, as well as books.…

I love the Diderot quote. I'd also encountered earlier:

"As long as the centuries continue to unfold, the number of books will grow continually, and one can predict that a time will come when it will be almost as difficult to learn anything from books as from the direct study of the whole universe. It will be almost as convenient to search for some bit of truth concealed in nature as it will be to find it hidden away in an immense multitude of bound volumes. When that time comes, a project, until then neglected because the need for it was not felt, will have to be undertaken...."

... and on for another several paragraphs. It's an extraordinarily keen observation on the state and future of knowledge. At the always excellent History of Information website:

http://www.historyofinformation.com/detail.php?entryid=2877

(Diderot is on my list of authors to explore in more depth.)

The fact that the quality of any given information or exchange is often (though not always) entirely divorced from its source (or author) is another interesting note. There are a few points here worth expanding on.

At least probabalistically, there are spaces (real or virtual) in which it's more likely to encounter good ideas. HN for its various failings, does well in today's Net. Google+, for all its faults, was similarly useful.

Size matters far less than selection. The tendency for centres of learning, research, and/or inquiry (and not necessarily in that order) to emerge is one that's been long observed, and their durability remarkable. The first universities (Bologna, Padua, Oxford, Paris, Cambridge, Heidelberg, and others, see: https://en.wikipedia.org/wiki/Medieval_university) are often still, 600 - 700 years later among the best in the world. Certainly in the US, Harvard, Yale, Princeton, M.I.T., among the earliest founded, remain the most prestigious. Though as noted in the conversation with Tyler Cowen and Patrick Collison, the list from 1920 is "completely the same, except we’ve added on California".

https://conversationswithtyler.com/episodes/mark-zuckerberg-...

What happens as the overal quantity and flux of information increases is that more effective rejection systems are required. That is: you've got too much information flowing in, you want a way to cheaply, with minimal effort or consequential residiual load, reject information that may be irrelevant, with minimal bias.

There are numerous systems that have been arrived at, and many of our cognitive biases or informal tests for truth arise out of these (optimism, pessimism, availability, sunk-cost, tradition, popularity, socio-ethnic prejudice, etc.). Randomised methods are probably far fairer and less prone to category error. Michael Schulson's sortition essay in Aeon remains among the best articles I've read in the past decade, if not several:

"If You Can't Choose Wisely, Choose Randomly"

https://aeon.co/essays/if-you-can-t-choose-wisely-choose-ran...

Another fundamental problem is self-dealing and self-selection within institutions. Much of the failure within academia (also touched on by Cowen and Collison, who, I'll note, I don't generally agree with, though they are touching on and making many points I've been pursuing for some years) comes from the fact that it's internal selection of students, faculty, articles, topics, and ideologies, rather than strict tests of real-world validity, which promote these structures.

The same problems infect government and business -- it's not as if any one social domain is immune to this.

Oh, and another lecture by H.G. Wells on that topic:

"...When I go to see my government in Westminster I find presiding over it the Speaker in a wig and a costume of the time of Dean Swift, the procedure is in its essence very much the same. The Members debate bring motions and when they divide the art of counting still in governing bodies being in its infancy they crowd into lobbies and are counted just as a drover would have counted his sheep two thousand years ago...."

https://invidio.us/watch?v=qRgP-46AC_o

(Audio quality is exceptionally poor, 1931 recording.)

Partial transcript: http://www.aparchive.com/metadata/INTERVIEW-WITH-H-G-WELLS-S...

AI ... may be useful, but seems to be result-without-explanation, a possible new form of knowledge, to go with revelation (pervasive if not particularly acurate), technical (means), and scientific (causes / structural).

Wholehearted agreement on LibGen.

Very enjoyable conversation BTW, thank you.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#257

Earlier quoted context omitted.

No one is proposing we use floppy disks. Redundant, shared servers ARE a forever solution. Making sure your data is one one of the ones that makes it seems like a vastly easier proposition to me than writing data to clay tablets and trying to keep those from ending up in a dump somewhere.

What is the likelihood that historians a century or two hence will have an application capable of turning an ISO 32000-1 file into a human-readable text? If we are talking about archaeologists, rather than historians, even ASCII and Unicode could be a challenge to work out.

Because those hundreds of years don't transpire in a glimpse. At some point in the middle there will be deprecated formats and new ones, and transcoders you can batch run. Sure it relies on intervention, but the upside is any/everyone else can copy the one persons work.

Yes we should learn from history, but we should also not assume that everything that happened before will happen the same way again, given how much of our world has changed.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#258

Earlier quoted context omitted.

While brushing up on the encyclopaedists, I found this little gem: "Among some excellent men, there were some weak, average, and absolutely bad ones. From this mixture in the publication, we find the draft of a schoolboy next to a masterpiece." — Denis Diderot Taking the quote out of context (and aside from its historical male-centered language) - it sure rings true of the current state of the web, as well as books.…

I love the Diderot quote. I'd also encountered earlier: "As long as the centuries continue to unfold, the number of books will grow continually, and one can predict that a time will come when it will be almost as difficult to learn anything from books as from the direct study of the whole universe. It will be almost as convenient to search for some bit of truth concealed in nature as it will be to find it hidden away…

Nature shows us how to process information at ever increasing noise and scale - https://www.edge.org/response-detail/10464

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#259
post #257

Earlier quoted context omitted.

What is the likelihood that historians a century or two hence will have an application capable of turning an ISO 32000-1 file into a human-readable text? If we are talking about archaeologists, rather than historians, even ASCII and Unicode could be a challenge to work out.

Because those hundreds of years don't transpire in a glimpse. At some point in the middle there will be deprecated formats and new ones, and transcoders you can batch run. Sure it relies on intervention, but the upside is any/everyone else can copy the one persons work. Yes we should learn from history, but we should also not assume that everything that happened before will happen the same way again, given how much o…

> However, without archivists actively transforming content to new formats as required, it might only take a few decades before a lot of content starts to require a massive effort to read.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#260
post #258

Earlier quoted context omitted.

I love the Diderot quote. I'd also encountered earlier: "As long as the centuries continue to unfold, the number of books will grow continually, and one can predict that a time will come when it will be almost as difficult to learn anything from books as from the direct study of the whole universe. It will be almost as convenient to search for some bit of truth concealed in nature as it will be to find it hidden away…

Nature shows us how to process information at ever increasing noise and scale - https://www.edge.org/response-detail/10464

Yes and no.

Briefly: the article distinguishes "endocrinal" vs. "distributed" decisionmaking.

This applies at some levels, but not at others.

For individual humans, we don't have the option of rewiring our concsiousnesses, which are rather pathetically single-threaded, and can at best multitask poorly by task-switching, at a very great loss of task proficiency.

Even withing collective organisations (companies, governments, organisations, communities), the multiple independent actors works where those actors' actions are autonomous and independent of others. Or, in the alternative, where they work without mutual conflict toward a common goal.

But you get problems where either individual actors' motivations and actions are in conflict, or in which a single global decision must be made (as with various global catastrophic risks), and multiple independent decisions cannot be arrived at. Even for noncritical arbitrary decisions, such as which side of the road to drive on, in which there is no compelling argument to be made for one side or the other, but in which both sides cannot be simultaneously selected, you need some global decisionmaking capacity.

When you reach the point of either an existing decisionmaking system (as in: a single human, with the finite and largely immutable information acquisition and processing capabilities corresponding), or a multi-agent system which must reach a common decision, you've got the challenge of limiting data intake to that amount which allows effective function within the environment, and avoids overloading capabilities or ineffective action.

Post reply on HN