Live data from Hacker News

Archivists Are Trying to Make Sure LibGen Never Goes Down

vice.com

71–80 of 270 posts

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#71
What's interesting is that 32TB is becoming more and more affordable and the research material is roughly staying about the same size.

That might change though as people start including video + data within papers and have new notebook formats that are live and contain docker containers/ipython, etc.

It's a shame we can't just mail these around.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#72
post #2

This is an extremely important effort. The LibGen archive contains around 32 TBs of books (by far the most common being scientific books and textbooks, with a healthy dose of non-STEM). The SciMag archive, backing up Sci-Hub, clocks in at around 67 TBs [0]. This is invaluable data that should not be lost. If you want to contribute, here's a few ways to do so. If you wish to donate bandwidth or storage, I personally k…

For important archives like this maybe we need some sort of turn-key solution for the masses? Like a Raspberry Pi image that maintains a partial mirror. Imagine if one could by a RPi and external HD, burn the image, and connect it to some random wifi network (at home, at work, at the library, etc).

I'm not hosting a copy of this at work (where we easily have 32TB on old hardware) since distributing it is copyright infringement. The same goes for my home connection.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#73

What's interesting is that 32TB is becoming more and more affordable and the research material is roughly staying about the same size. That might change though as people start including video + data within papers and have new notebook formats that are live and contain docker containers/ipython, etc. It's a shame we can't just mail these around.

[deleted]

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#74
post #60

Earlier quoted context omitted.

Do we have anything rated for a few millennia of interstellar radiation besides etched gold plates?

whats wrong with etched gold disks?

I wonder if that's actually feasible for this application.

Microdots managed about 32MiB / square centimeter (from the "you could fit the Bible 50 times in a square inch") measure. That's a completely arbitrary density to achieve, since it was photographs which were enlarged and shrunk, and you could hypothetically use any encoding for your gold etchings, and it also leaves out the question of "how fine can you etch gold plates while still having the engraving be 'robust'".

But in any case, that gives you a target area of ~325 square meters for a 100TiB archive.

That's a lot, but not like a crazy obviously impossible number like a million square km or something.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#76
post #19

Why not publish the site over IPFS, that would make P2P hosting much simpler?

Currently (at least for the-eye) it's about IPFS's barrier of entry. I expect LibGen's case to be similar. Most people don't know about it, and if even those that knew about it had to learn how IPFS works etc, they would probably just try to find the book they're looking for elsewhere.

No need to conflate the frontend (the end-user interface that 'most people' use when trying to 'find the book they're looking for') with the mirroring/archiving backend (the distributed/p2p technology used to 'make sure LibGen never goes down').

The frontend would still be a user-friendly HTTP web-application (or collection of several) that pulls (portions of) the archive from the distributed/resilient backend to serve individual files to clients.

The backend can be a relatively obscure, geeky, post-BitTorrent p2p software like IPFS or Dat, as long as those willing to donate bandwidth/storage can run it on their systems. This is a vastly different audience from 'most people'.

The real question is which software's features best fits the backend use-case (efficiently hosting a very large and growing/evolving, IP-infringing dataset). Dat [1] has features to (1) update data and efficiently synchronize changes, and to (2) efficiently provide random-access data from larger datasets. Two quite compelling advancements over BitTorrent for this use-case.

[1] https://docs.datproject.org/docs/faq#how-is-dat-different-th...

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#77
post #11
post #7

Is there a way to just download the whole 32TB to your own machine? I see a ton of mirrors but the content seems to be highly fragmented between them

There are ways to do so. The archive is made up of many, many torrents (I believe it's a monthly if not biweekly update of the database). If you have the storage/bandwidth availability for the whole 32TBs, please get in touch and I may be able to help you get the whole deal without too much hassle. Otherwise, just pick some torrents (it would be best to pick them based on torrent health, but they are so many to check…

Why doesn't someone maintain a single torrent containing a snapshot of the full archive at a given point in time, updated (say) monthly?

I want a full mirror, and ain't nobody got time to deal with 2000 torrents, many of which have no seeders. That's a really dumb way to run this particular railroad.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#79
post #60

Earlier quoted context omitted.

whats wrong with etched gold disks?

I wonder if that's actually feasible for this application. Microdots managed about 32MiB / square centimeter (from the "you could fit the Bible 50 times in a square inch") measure. That's a completely arbitrary density to achieve, since it was photographs which were enlarged and shrunk, and you could hypothetically use any encoding for your gold etchings, and it also leaves out the question of "how fine can you etch…

Assuming you can cut the etching down to 1m squares and stack them on top of one another that should be no problem at all. Assuming each layer is 1.2mm thick (same as a CD) that's a volume of 1m x 1m x .390m, it would fit easily in a payload fairing of any rocket that can get it into orbit.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#80
post #58

Earlier quoted context omitted.

From what I've heard a good chunk of people rotate their seeds for LibGen because their seedboxes can't handle all the connections for every torrent at once.

Is there some tool or documentation describing this practice?

I'm sure someone could get you the info to get setup as a seeder. For modern clients it's rather rather trivial to manage that many torrents. Get any decent modern CPU, 4gb+ ram, and $560 in storage and you're off.
Post reply on HN