Live data from Hacker News

Archivists Are Trying to Make Sure LibGen Never Goes Down

vice.com

231–240 of 270 posts

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#231
post #215

Earlier quoted context omitted.

No, that's BS. There are well known cases of genetics and cybernetics being banned for ideological reasons during Stalin's time. Scientific books and articles of convicted 'enemies of the state' were dangerous to possess in that time too. Some scientists used ideological 'arguments' in scientific debates which were dangerous to argue against. But all that, AFAIK, ended after Stalin's death in 1953. Moreover, I've nev…

Not sure what you are saying. Mathematicians were not even allowed to travel abroad [1] and any "concessions" were essentially as it pleased the USSR state. Only from 1990 was movement free in the true sense of the word. [1] An example was when Margulis won the Fields medal: https://en.wikipedia.org/wiki/Grigory_Margulis . There are many other examples too.

What does that have to do with sharing knowledge in the USSR and the countries in the Soviet block?

It was never in Soviet ideology to hide knowledge behind paywalls. See, for example, this [0] post about Mir publishing house and warm comments of Indians who grew up with their books. Sci-hub's ideology is just continuation of this approach.

[0] https://news.ycombinator.com/item?id=21352277

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#232

Earlier quoted context omitted.

> possibly beaten only by Sci-Hub Today I learned that Library Genesis is actually "powered by Sci-Hub" as its primary source. So I guess they're sister projects by similarly minded people (who seem to be mostly/originally based in Slavic countries, which I find interesting culturally - perhaps it's due to a looser legal environment + activist academics?). > Just about everybody in academia knows about it. That reall…

who seem to be mostly/originally based in Slavic countries, which I find interesting culturally - perhaps it's due to a looser legal environment + activist academics? You see the same situation with Asia --- it's a collectivist culture, they have a very different perspective on IP in general.

That makes sense, thanks for pointing that out. I can see how it's related to the value system of collectivist (as opposed to individualist) cultures, and how they see intellectual property - as a common good, beyond personal/private ownership.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#233

Earlier quoted context omitted.

This is not a solvable problem without technological continuity, or some unimaginably smart technology we can't imagine today. If you found a mysterious archive object and had no idea what it was - CD-R, hard drive, SSD, whatever - not only would you have to reinvent an entire hardware reader around it, you would also have to work out the file structure, extract the data (some of which could be damaged), and reverse…

Take a CD-R of some MP3 with English language file names stored on a FAT32 filesystem for example. Assume the reflective layer didn't rust since it was abandoned in a dry climate and our future archaeologist has access to roughly modern levels of technology. 1. Even if the CD-R has been crushed and shattered you could use a modern and cheap microscope to read continuous pits and lands off the disk [0,1]. It would be…

I mean, archaeology and linguistics have been figuring out ancient languages as an entire field, while determined individual hobbyists are able to reverse engineer unknown file formats.

By which I mean, many file formats are syntactically much simpler and more obviously structured than natural languages. It might take an entire field to reverse engineer weird formats like .DOC once all knowledge gets lost, but I doubt this will be the case for bitmaps or UTF-8 ...

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#234

Earlier quoted context omitted.

> Somewhere at Google there is a database containing 25 million books and nobody is allowed to read them. Indeed, what an intellectual tragedy.. > In August 2010, Google put out a blog post announcing that there were 129,864,880 books in the world. The company said they were going to scan them all. That seems like a surprisingly "small" number. Well, in trying to picture a physical library with 130 million books, may…

Until fairly recently (historically), books were overwhelmingly scarce. A few datapoints: - The total number of books -- not titles, but actual bound volumes -- in Europe as of 1500 CE, was about 50,000. By 1800, the total was just under one billion. - The library of the University of Paris circa 1000 CE comprised about 2,000 volumes. It was among the largest in Europe. - The Library of Constantinople in the 5th cent…

Thank you for that, very interesting and educational. I love how you led up to the punchline. It made me see that books as a technology and artifact are part of the "history of information", and how books are becoming subsumed in a shared trajectory with media/data in general.

> half of all the recorded information of humankind was created in the past two years

That is shocking to imagine, and it's exponentially growing.

It reminds me of Vannevar Bush's "As We May Think", pointing out the emerging information overload in society. It certainly puts things in perspective, how we (humanity) have been making a conscious, collaborative effort to develop globally networked computers, one of whose important functions is to help us organize all the information, including books.

The conundrum it seems is that technology is also a massive multiplier/amplifier of the amount of data, that its capacity to help us organize would never catch up to what it's helping to produce.

> total storage for the 38 million volumes of the Library of Congress would be slightly under 200 TB

I guess it's redundant to say, but I'm sure in the near future that would fit on a thumb drive!

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#235
post #130

Libgen is one of the greatest contributors to scientific productivity worldwide, possibly beaten only by Sci-Hub. Just about everybody in academia knows about it. If it ever vanished, some of us could probably still get by trading files from person to person, but nothing could be as perfect as what we got now.

> Just about everybody in academia knows about it Just about everybody in academia uses it , too, especially in the case of Scihub. I can't imagine taking the time to actually check whether I have access to some journal when I want to read a paper, let alone jump through all the hoops before you can get a PDF. The first thing we did when my partner's paper was recently published was check to see if it was on Scihub y…

I remember, in the early 2000s, going through all the trouble of logging into my university library's proxy portal to get access to certain scientific papers. I probably wouldn't have done that if SciHub was available, and it probably would have opened up my eyes sooner to the fact that most people don't even have access to such a portal. Although, frankly, it was a different web back then and if you were persistent you could actually find anything.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#236

Earlier quoted context omitted.

> Lastly, you can always contribute books. If you buy a textbook or book, consider uploading it (and scanning it, should it be a physical book) in case it isn't already present in the database. There's no easy solution for scanning physical books, is there?

There are providers [1] that will destructively scan the book for you and return a PDF. If you want to preserve the book, you're stuck using a scanning rig [2]. The Internet Archive will also non-destructively scan as part of Open Library [3], but they only permit one checkout at a time of scanned works, and the latency can be high between sending them a book and it becoming available. FYI, 600 DPI is preferred for a…

[deleted]

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#237

Earlier quoted context omitted.

It’s easy to take this stance in a rich country. But what about the people in countries where one of these books cost the equivalent of a year’s wages. Not so black and white eh?

As far as I know, prices of books differ between rich and developing nations. For e.g., The C Programming Language that costs $50 in the US [1], is sold for Rs. 259 (~$4 US) in India. I believe that is the case with most "economy editions" specifically targeted at developing nations. It certainly isn't an "year's wages". While I do understand your point, it still does not justify encouraging modern-day Robinhoods' an…

Economy editions are nice, but only make up a tiny fraction of what's out there to consume. Again, $4 is great for a relatively rich nation like India, but what about Eritrea?

Maybe that was a little flippant, but The Law in a big rich country is pretty meaningless to someone trying to make a better life for themselves in a poor country.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#238

Earlier quoted context omitted.

> It's absolutely unkillable Just like any other distributed system, this is vulnerable to organized take downs and scare tactics. There was a whole bunch of mirrors of Pirate Bay, yet once most of Europe's legal systems adopted the "sharing is theft" mindset, it became pretty much impossible to find one.

I was just in Europe and used piratebay while I was there. The main site didn’t work, but searching “piratebay mirror” found one that did right away.

Here in Norway, ISPs are actually legally obligated to block access to The Pirate Bay. Mirrors work.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#239

Earlier quoted context omitted.

I'm incredibly thankful that the public library system was invented before copyright maximalists got control of Congress.

In this case the issue seems to have come from "copyright minimalists" instead : wanting the books to be freely available, rather than making money for Google... I wonder why the Copyright Office didn't just buy Google Books, would only have cost a few hundred million $ ?

> Upon hearing that Google was taking millions of books out of libraries, scanning them, and returning them as if nothing had happened, authors and publishers filed suit against the company, alleging, as the authors put it simply in their initial complaint, “massive copyright infringement.”

This is where the project derailed and never quite recovered.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#240

Earlier quoted context omitted.

Take a CD-R of some MP3 with English language file names stored on a FAT32 filesystem for example. Assume the reflective layer didn't rust since it was abandoned in a dry climate and our future archaeologist has access to roughly modern levels of technology. 1. Even if the CD-R has been crushed and shattered you could use a modern and cheap microscope to read continuous pits and lands off the disk [0,1]. It would be…

I mean, archaeology and linguistics have been figuring out ancient languages as an entire field , while determined individual hobbyists are able to reverse engineer unknown file formats. By which I mean, many file formats are syntactically much simpler and more obviously structured than natural languages. It might take an entire field to reverse engineer weird formats like .DOC once all knowledge gets lost, but I dou…

Bitmaps are easy enough, but I wouldn't bet on UTF-8.

And any modern compression is probably right out without technological continuity.

Post reply on HN