Live data from Hacker News

Archivists Are Trying to Make Sure LibGen Never Goes Down

vice.com

121–130 of 270 posts

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#121

Earlier quoted context omitted.

It is very common these days to buy the WD 8TB, 10TB and 12TB external USB3 hard drives and remove their cases, and put them in some sort of home built file server or NAS. There's a technique to put a thin section of kapton tape on one of the SATA pins so that they will power up from ordinary PC/ATX type power supplies with regular SATA power connectors. https://www.instructables.com/id/How-to-Fix-the-33V-Pin-Issu...…

> at no greater or lesser annual failure rate than the expensive enterprise hard drives. I've read these reports as well, but I can say that it's not my experience (we've gone through a few rounds of shucking at the Internet Archive, for economy and in one case necessity after the 2011 Thailand floods pinched the supply chain). Our raw failure rates on shucked drives are significantly higher, and the drives themselve…

From reddit.com/r/datahoarder I don't believe I've seen a single instance of 8, 10 or 12TB consumer USB3 drives coming out of the plastic case as a model that is shingled recording. The average consumer trying to copy many dozens of GB onto an external drive would not tolerate SMR write performance.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#122
post #2

This is an extremely important effort. The LibGen archive contains around 32 TBs of books (by far the most common being scientific books and textbooks, with a healthy dose of non-STEM). The SciMag archive, backing up Sci-Hub, clocks in at around 67 TBs [0]. This is invaluable data that should not be lost. If you want to contribute, here's a few ways to do so. If you wish to donate bandwidth or storage, I personally k…

> Lastly, you can always contribute books. If you buy a textbook or book, consider uploading it (and scanning it, should it be a physical book) in case it isn't already present in the database. There's no easy solution for scanning physical books, is there?

Your local physical library may make a book scanner available. Mine does, with a posted 60-pages-at-a-time limit (though I don't know how this is enforced).

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#123
post #5

one of the next interplanetary or Interstellar Probe should carry a copy of the sci-hub torrent in some kind of permanent storage

there is no need to put a data storage archive on something shot out into interstellar space. geostationary telecommunications satellites are at a sufficiently high enough orbit that they will likely outlast human civilization. We could destroy ourselves with nuclear war, regress to a stone age level of technology, rediscover spaceflight and go find them long before the orbits of any of them decay.

too close to the warzone to survive

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#124

Earlier quoted context omitted.

I noticed this yesterday while shopping for cyber Monday deals. If you want to load up a server with drives, perhaps the external drives can be removed from their cases and used internally?

Check it before you buy it. Years ago I bought a 1TB WD external drive. The usb interface was connected directly to the drive.

Thanks! I will have to read up on shucking and and /r/DataHoarder. I would think someone already has a list of which external drives can be used this way.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#125
post #11

Earlier quoted context omitted.

There are ways to do so. The archive is made up of many, many torrents (I believe it's a monthly if not biweekly update of the database). If you have the storage/bandwidth availability for the whole 32TBs, please get in touch and I may be able to help you get the whole deal without too much hassle. Otherwise, just pick some torrents (it would be best to pick them based on torrent health, but they are so many to check…

Why doesn't someone maintain a single torrent containing a snapshot of the full archive at a given point in time, updated (say) monthly? I want a full mirror, and ain't nobody got time to deal with 2000 torrents, many of which have no seeders. That's a really dumb way to run this particular railroad.

Because torrent clients can't handle that many pieces in a single torrent. There are algorithms that are super-linear, maybe even quadratic or worse. They start causing trouble int eh TB range.

Also the UI for adding many torrents is much nicer than for selecting a non-trivial subset of files inside a single torrent. Also many parts of the ecosystem handle partial-seeds that do and will only for the near future seed a subset and not leech any other parts. They often get treated as leechers, despite not really being leechers.

TL;DR: 2k files are just a watchfolder and a cp * watchfolder/ away from working. Scaling does not work with one fat 32TB, however.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#126

Earlier quoted context omitted.

there is no need to put a data storage archive on something shot out into interstellar space. geostationary telecommunications satellites are at a sufficiently high enough orbit that they will likely outlast human civilization. We could destroy ourselves with nuclear war, regress to a stone age level of technology, rediscover spaceflight and go find them long before the orbits of any of them decay.

too close to the warzone to survive

okay, high Mars orbit. Quite a lot less delta-v requirement for some theoretical several thousand kilogram chunk of long-lifespan data storage, and easier to discover and retrieve, than achieving solar escape velocity (look at the delta-v budget for the new horizons probe vs. its total size and weight, for instance).

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#127
post #40

Earlier quoted context omitted.

It’s incredibly sad: https://www.theatlantic.com/technology/archive/2017/04/the-t...

I'm incredibly thankful that the public library system was invented before copyright maximalists got control of Congress.

In this case the issue seems to have come from "copyright minimalists" instead : wanting the books to be freely available, rather than making money for Google...

I wonder why the Copyright Office didn't just buy Google Books, would only have cost a few hundred million $ ?

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#129

What's interesting is that 32TB is becoming more and more affordable and the research material is roughly staying about the same size. That might change though as people start including video + data within papers and have new notebook formats that are live and contain docker containers/ipython, etc. It's a shame we can't just mail these around.

You can buy 48TB (4x12TB) for €1000. Store some index on an SSD, and you have another full node.

If you don't care about warranty, 8 and 12TB drives routinely go for $15/TB on sale inside WD Elements.

I picked up 32TB for just under $500 with discount over the holiday that way.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#130
Libgen is one of the greatest contributors to scientific productivity worldwide, possibly beaten only by Sci-Hub. Just about everybody in academia knows about it. If it ever vanished, some of us could probably still get by trading files from person to person, but nothing could be as perfect as what we got now.
Post reply on HN