Live data from Hacker News

Archivists Are Trying to Make Sure LibGen Never Goes Down

vice.com

241–250 of 270 posts

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#241

Earlier quoted context omitted.

In this case the issue seems to have come from "copyright minimalists" instead : wanting the books to be freely available, rather than making money for Google... I wonder why the Copyright Office didn't just buy Google Books, would only have cost a few hundred million $ ?

> Upon hearing that Google was taking millions of books out of libraries, scanning them, and returning them as if nothing had happened, authors and publishers filed suit against the company, alleging, as the authors put it simply in their initial complaint, “massive copyright infringement.” This is where the project derailed and never quite recovered.

Did we read the same article ?

EDIT :

> As Tim Wu pointed out in a 2003 law review article, what usually becomes of these battles—what happened with piano rolls, with records, with radio, and with cable—isn’t that copyright holders squash the new technology. Instead, they cut a deal and start making money from it.

[...]

> now, in 2011, there was a plan—a plan that seemed to work equally well for everyone at the table

[...]

> DOJ’s intervention likely spelled the end of the settlement agreement. No one is quite sure why the DOJ decided to take a stand instead of remaining neutral. Dan Clancy, the Google engineering lead on the project who helped design the settlement, thinks that it was a particular brand of objector—not Google’s competitors but “sympathetic entities” you’d think would be in favor of it, like library enthusiasts, academic authors, and so on—that ultimately flipped the DOJ.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#242

Earlier quoted context omitted.

Indeed, but for how long, considering the Google Books team itself seemingly wants to delete the ~100M books database ? P.S.: Might be soon, considering that it was basically what Google was initially about, and one of the founders just resigned from Alphabet...

Source on the deletion initiative?

https://news.ycombinator.com/item?id=21694061

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#243

Earlier quoted context omitted.

Most people don't care. The chance anything at all bad will happen is so incredibly low.

This isn't even movies wherein some large studio's can send notices. I don't think publishing houses have that many funds to send so many legal notices. Books are a safe bet to pirate

At least the large academic publishers are sitting on enormous stacks of cash, so that argument doesn't fly.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#245

Earlier quoted context omitted.

Until fairly recently (historically), books were overwhelmingly scarce. A few datapoints: - The total number of books -- not titles, but actual bound volumes -- in Europe as of 1500 CE, was about 50,000. By 1800, the total was just under one billion. - The library of the University of Paris circa 1000 CE comprised about 2,000 volumes. It was among the largest in Europe. - The Library of Constantinople in the 5th cent…

Thank you for that, very interesting and educational. I love how you led up to the punchline. It made me see that books as a technology and artifact are part of the "history of information", and how books are becoming subsumed in a shared trajectory with media/data in general. > half of all the recorded information of humankind was created in the past two years That is shocking to imagine, and it's exponentially grow…

Bush's essay is of course a classic. There are some precursors -- there's a BBC interview of H.G. Wells describing something similar from the 1940s.[1] E.F. Forster's The Machine Stops has some similar ideas. And various encyclopaedists very much embodied similar ideals.

I've been listening to Peter Adamson's "History of Philsophy Without Any Gaps" podcast, which is excellent, and spends a fair bit of time looking at the historiography of the topic -- what works were preserved, how, various interpretations, practices, preservation, and losses. Interesting to note that most of the preserved Greek and Roman works were found in obscure Arabian monastaries and libraries. The mainstream collections themselves were often lost in raids, fires, or other mishaps. Which makes the LibGen situation all the more relevant and urgent.

(I'm a huge user of the site and others like it, for what it's worth.)

On the amount of total data being captured: there's a huge difference between quantity and quality measures of information. They're almost certainly inversely related.

Of what books were written in antiquity, up to the time of the printing press, say, odds were fairly strong that a work would be read.

At 1 million new titles being published per year, there are only 330 people in the US per book, or roughly 400 native English speakers worldwide. (With ~2 billion speakers worldwide, the total audience might reach 2,000 per book). Clearly, most of what's being written will have a very small, or no, audience.

For machine-captured data, the likelihood that any of it is seen directly by a human is vanishingly small. More of it will undergo some level of machine processing or interpretation, though even that only applies to a fairly small fraction of data. Insert old joke about the WORN drive: write once, read never.

As for storage costs (and/or size), at a 15% cost reduction per year, storage halves every 4.67 years (4 years and 8 months), which means that in 10 years, the $10k price tag becomes $2k, and in 20 years, it should be under $400. For the entire Library of Congress collection.

Flash drives seem to be increasing in capacity by a factor of 10 every 2.5 years. There are now 2 TB flash drives, so 200 TB might be as little as 5 years out. That ... still sounds optimistic to me.

https://m.eet.com/media/1171702/digital_storage_in_consumer_...

https://www.digitaltrends.com/computing/largest-flash-drives...

The more practical problems are simply organising, cataloguing, and accessing the archives. This is an area that still needs help.

________________________________

Notes:

1. I think that's from "Science and the Citizen*, 1943, though the BBC and I have a disagreement concerning access. https://www.bbc.co.uk/archive/hg-wells--science-and-the-citi...

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#246
post #225
post #100

Earlier quoted context omitted.

An idea I've seen is including messages at several levels. At the outermost level you describe in very basic format how to build a magnifying glass. From there you have diagrams that are legible that describes how to build a microscope. From there you have more than enough space to describe the basics of what else is in there and to start describing your language. I'm thinking optical storage in a clear rock of some…

>I'm not trying to spec "Ugh wanders out of the jungle, sees our pretty rock, and personally has a 20th century civilization up and running in 10 years" or anything crazy. Quite coincidental, there is currently an Anime running named "Dr. Stone" which is quite exactly about that; jump starting human civilizatio from the stone age to modern day as fast as possible. Atleast in-story it's been a few months and they're c…

Yeah, it's on my list. My strategy is generally to wait until I can just mainline the entire season, so I tend not to watch the latest seasonals, but it's definitely on my list when the season is done.

In Vernor Vinge's "A Fire Upon the Deep", a very advanced civilization in the outer galaxy that can't reach where we are for $REASONS has as a persistent hobby speculation on the fastest way to bootstrap advanced civilizations, assuming essentially-perfect knowledge of physics instead of blundering around.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#247

Maybe we should print this out on acid-free paper-thin flexible wood-pulp sheets stitched to together to form linear organized aggregations. Each aggregation would contain one or more works and be searchable using a SQL-like database. To make this plan really work there would need to be a collection of geographically distributed long term physical repositories that would receive periodic updates as new material becam…

Forget DRM, even future Engilsh may be incomprehensible. There is an entire field of study dedicated to finding a way to make our future voice heard, without a good plausible solution, called Nuclear Semiotics (https://en.wikipedia.org/wiki/Nuclear_semiotics).

If we can't effectively warn a future (>10,000 years) generation to stay away from something that may harm or kill them, what chance do we have of making a universally understandable archive of data?

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#248

Earlier quoted context omitted.

I mean, archaeology and linguistics have been figuring out ancient languages as an entire field , while determined individual hobbyists are able to reverse engineer unknown file formats. By which I mean, many file formats are syntactically much simpler and more obviously structured than natural languages. It might take an entire field to reverse engineer weird formats like .DOC once all knowledge gets lost, but I dou…

Bitmaps are easy enough, but I wouldn't bet on UTF-8. And any modern compression is probably right out without technological continuity.

I think if you gave a philologist living in 1880 AD a clay tablet with a binary inscription of a fragment of an English poem encoded UTF-8 they would decode it very quickly.

This is what the philologist would see:

>...ABABABBBABBABAAAABBBBAABAABABBAAAABAAAAAABBABAABABBAABBAAABAAAAAAABAABBBABBBABAAABBABAABABBBAABBAABAAAAAABBAABAAABBAAAABABBABBBAABBAAABBABBABAABABBABBBAABBAABBBAABAAAAAABBBBAABABBABBBBABBBABABAABAAAAAABBBABBBABBABBBBABBBABABABBABBAAABBAABAAAABAAAAAABBAAABAABBAABABAABABBAAAAAABABAABABABAAABBABAAAABBAABABABBBAABAABBAABABAABAABBBABBBAABBAABAAAAAABBAAABAABBBAABAABBABAABABBBAABBABBABABBABBAABABABBBAABAAABAAAAAABBBAAAAABBABAABABBBAAAAABB...

How it would probably go:

1. Hmmmm there are only two symbols A and B, these symbols can't be words since no language has only two words. Thus the words must be made of a string of these symbols.

2. Every 8-th symbol* is a A. Lets try putting the symbols in groups of size 8.

3. These groups of 8 can't be words because they repeat far too often and they would only allow 128 possible words. Thus these groups of 8 might be letters in an alphabet.

4. Does the frequency of this possible letters fit any known languages? Yes, English.

5. Which group of 8 is "e"?

A few minutes later and the clay tablet is decoded.

* - This is not always true in utf-8 but true in most encoding of Latin alphabets including this example. Even with some variable length characters thrown in this fact would stand out.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#249

Earlier quoted context omitted.

Thank you for that, very interesting and educational. I love how you led up to the punchline. It made me see that books as a technology and artifact are part of the "history of information", and how books are becoming subsumed in a shared trajectory with media/data in general. > half of all the recorded information of humankind was created in the past two years That is shocking to imagine, and it's exponentially grow…

Bush's essay is of course a classic. There are some precursors -- there's a BBC interview of H.G. Wells describing something similar from the 1940s.[1] E.F. Forster's The Machine Stops has some similar ideas. And various encyclopaedists very much embodied similar ideals. I've been listening to Peter Adamson's "History of Philsophy Without Any Gaps" podcast, which is excellent , and spends a fair bit of time looking a…

While brushing up on the encyclopaedists, I found this little gem:

"Among some excellent men, there were some weak, average, and absolutely bad ones. From this mixture in the publication, we find the draft of a schoolboy next to a masterpiece." — Denis Diderot

Taking the quote out of context (and aside from its historical male-centered language) - it sure rings true of the current state of the web, as well as books.

About the inverse relationship of quantity vs quality, we seem to be drowning in quantity! As you've pointed out, there's great need for thoughtful organization and curation.

I like how you break down the quantifiable aspects to draw a historical trend and future projection. The rise of "data science" and "big data" in the past few decades really makes sense in this light.

I'm sure machine learning and "AI" will play an increasing role in the task of organizing and processing all this information, but at the bottom I feel that the most value probably comes from human curation.

LibGen has been an amazing resource for me as a lover of knowledge, a life-long book worm. I've got bookshelves and boxes full of physical books as well, but it's a drop in the ocean..

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#250
post #231

Earlier quoted context omitted.

Not sure what you are saying. Mathematicians were not even allowed to travel abroad [1] and any "concessions" were essentially as it pleased the USSR state. Only from 1990 was movement free in the true sense of the word. [1] An example was when Margulis won the Fields medal: https://en.wikipedia.org/wiki/Grigory_Margulis . There are many other examples too.

What does that have to do with sharing knowledge in the USSR and the countries in the Soviet block? It was never in Soviet ideology to hide knowledge behind paywalls. See, for example, this [0] post about Mir publishing house and warm comments of Indians who grew up with their books. Sci-hub's ideology is just continuation of this approach. [0] https://news.ycombinator.com/item?id=21352277

> Sci-hub's ideology is just continuation of this approach.

Actually, that was the point of what I was saying—the mathematicians had to be inventive and thus passed around preprints that they knew would also be read in the West.

Post reply on HN