Live data from Hacker News

Archivists Are Trying to Make Sure LibGen Never Goes Down

vice.com

111–120 of 270 posts

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#111

I don't see anyone having mentioned the possibility of posting this data to Usenet at all - at minimum for archival purposes which should be good for ~8-9 years. That way at least the data isn't lost. With so many of those torrents have 0 or 1 seed, this is a serious risk I think, despite the comments elsewhere about people rotating what they seed. I realize that doesn't solve the access problem for most people as mo…

Two thoughts on that. Encoding it to a text format with CRC data for posting to usenet is highly inefficient in terms of data storage. And 33TB of stuff is not going to be retained for 8-9 years, the last I checked due to the huge volume of binaries traffic, the major commercial usenet feed providers have at most 6-9 months of retention for the major binary groups. Beyond that it becomes cost prohibitive for them in…

Entirely agree about the lack of efficiency. No question about that.

However, in my personal experience, I have seen no issues downloading old data from any binary group. At least not with the provider I have. In fact, just this past week I obtained something sizable (several GBs) with no damaged parts so didn't even need the parchive recovery files at all. This has always been my experience. I've never seen anything like the pruning you are talking about. That sounds more like an issue with your specific provider to me.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#112

Earlier quoted context omitted.

> Lastly, you can always contribute books. If you buy a textbook or book, consider uploading it (and scanning it, should it be a physical book) in case it isn't already present in the database. There's no easy solution for scanning physical books, is there?

There are DIY book scanners ( http://diybookscanner.org ) and products such as the Fujitsu ScanSnap SV600. The SV600 has decent features like page-detection and finger-removal (I recommend using a pencil's eraser tip). I have personally used it to scan dozens of books, with satisfactory results.

Just saw a father who had to do it fully manually for her blind daughter. I shall show your comment to him.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#113
post #57

Yongle Encyclopedia was a similar project of the 15th century China. It was the largest encyclopedia in the world for 600 years until surpassed by Wikipedia. Alas, Yongle Encyclopedia is almost completely lost now. Archiving is harder than you think. https://en.wikipedia.org/wiki/Yongle_Encyclopedia

I read the Wikipedia article about it and the sad thing is that the majority of the Yongle Encyclopedia seem to have been destroyed only in quite recent times.

> but 90 percent of the 1567 manuscript survived until the Second Opium War in the Qing dynasty. In 1860, the Anglo-French invasion of Beijing resulted in extensive burning and looting of the city,[16] with the British and French soldiers taking large portions of the manuscript as souvenirs.

Preservation is easy if you don't get invaded.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#115
post #88
post #2

This is an extremely important effort. The LibGen archive contains around 32 TBs of books (by far the most common being scientific books and textbooks, with a healthy dose of non-STEM). The SciMag archive, backing up Sci-Hub, clocks in at around 67 TBs [0]. This is invaluable data that should not be lost. If you want to contribute, here's a few ways to do so. If you wish to donate bandwidth or storage, I personally k…

Sounds like anyone with a seed box could donate some bandwidth and storage by leeching then seeding part of it? It would be nice if there’s a list of seeder/leecher counts (like TPB) or better yet of priority list of parts that need more seeders. Edit: Found the other comment where you link to the seeding stats: https://docs.google.com/spreadsheets/d/1hqT7dVe8u09eatT93V2x...

Or better yet, a RSS feed that plays nice with auto-retention and quota settings. It just delivers you a bunch of parts that are in need of seeders and you use your existing mechanism to help with it.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#117
post #25

Related: Looking at harddisk cost per terabyte, quite often extern drives are cheaper than internal ones. For example right now in Germany I can get a WD 8TB USB 3.0 drive for 135€ but the cheapest internal 8TB drive costs 169€. Any idea why? It's puzzling.

It is very common these days to buy the WD 8TB, 10TB and 12TB external USB3 hard drives and remove their cases, and put them in some sort of home built file server or NAS. There's a technique to put a thin section of kapton tape on one of the SATA pins so that they will power up from ordinary PC/ATX type power supplies with regular SATA power connectors. https://www.instructables.com/id/How-to-Fix-the-33V-Pin-Issu...…

> at no greater or lesser annual failure rate than the expensive enterprise hard drives.

I've read these reports as well, but I can say that it's not my experience (we've gone through a few rounds of shucking at the Internet Archive, for economy and in one case necessity after the 2011 Thailand floods pinched the supply chain). Our raw failure rates on shucked drives are significantly higher, and the drives themselves are typically non-performant for high-throughput workloads (often being SMR disks/etc, though hopefully the move away from drive-managed SMR will finally kill that product category off).

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#118

Earlier quoted context omitted.

I'm sure someone could get you the info to get setup as a seeder. For modern clients it's rather rather trivial to manage that many torrents. Get any decent modern CPU, 4gb+ ram, and $560 in storage and you're off.

I think the problem is that because of the size of each torrent, and there's 1000 of them, it's difficult to effectively seed all at once, so instead people would rather seed sections at once, and rotate through them. I'm not sure how people setup the rotation though, that can't be an incredibly common feature but I could be wrong.

There are features that prioritize those with low seed/leech ratio in a sort of periodic fashion. Also it partially auto-balances because a swarm only needs a little more than unity ratio injected into it to get itself fully replicated. So each one that get's chosen because of a low seed/leech ratio will inherently drop out of that criteria as soon as the swarm is self-sufficient.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#119
post #6
post #3

There is a huge amount of duplication there (i.e. books that have many scans), I wonder if it would be better to tackle that versus doing a straight backup.

There are groups behind data curation as well, though it is much harder. LibGen sees an addition rate of about 230 GBs per month, while SciMag's is around 1.10 TBs per month. We should expect those numbers to increase in the future. The man-hours required to curate those database may very well cost much more than the storage and bandwidth required to store duplicates and incorrectly tagged files. In any case, as I sa…

Do you know if they process PDF to reduce file size ?
Post reply on HN