Live data from Hacker News

Archivists Are Trying to Make Sure LibGen Never Goes Down

vice.com

101–110 of 270 posts

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#101

Earlier quoted context omitted.

> Lastly, you can always contribute books. If you buy a textbook or book, consider uploading it (and scanning it, should it be a physical book) in case it isn't already present in the database. There's no easy solution for scanning physical books, is there?

Scanning with your phone is getting easier. At a minimum you can take a pic of each of the pages. Software can clean up the images, sorta. It's not ideal but it's better than nothing.

I found vFlat to be magical in cleaning up book scan images you took with your phone.

https://play.google.com/store/apps/details?id=com.voyagerx.s...

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#102
post #92
post #70

Earlier quoted context omitted.

This is my experience as well. In theory, IPFS is exactly the right thing for LibGen, but in practice I consider it unusable.

FWIW: StavrosK has actually been putting some serious effort into making IPFS accessible. See here for example: https://news.ycombinator.com/item?id=16521385

It's not really an accessibility issue so much as a performance one. If it was a reasonable alternative to something like the rsync daemon I'd use it all the time.

Unfortunately the performance issues and overhead is just too much.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#103
post #5

one of the next interplanetary or Interstellar Probe should carry a copy of the sci-hub torrent in some kind of permanent storage

This is an interesting idea because there's a lot of radical political and philosophical publications on there. Brill's Historical Materialsim book series is on there almost in its entirety.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#104
I don't see anyone having mentioned the possibility of posting this data to Usenet at all - at minimum for archival purposes which should be good for ~8-9 years. That way at least the data isn't lost. With so many of those torrents have 0 or 1 seed, this is a serious risk I think, despite the comments elsewhere about people rotating what they seed.

I realize that doesn't solve the access problem for most people as most of the users who need this research might not know how to use usenet or even be familiar with it at all, but I think the first major concern would be to secure the entire repository on a stable network. Usenet seems like a good place for that even if it doesn't serves as a means of distribution. Encrypting the uploads would make them immune to DMCA takedowns provided that the decryption keys weren't made public and were only shared with individuals related to the maintenance of the LibGen project.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#105

What's interesting is that 32TB is becoming more and more affordable and the research material is roughly staying about the same size. That might change though as people start including video + data within papers and have new notebook formats that are live and contain docker containers/ipython, etc. It's a shame we can't just mail these around.

You can buy 48TB (4x12TB) for €1000. Store some index on an SSD, and you have another full node.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#106
post #92

Earlier quoted context omitted.

FWIW: StavrosK has actually been putting some serious effort into making IPFS accessible. See here for example: https://news.ycombinator.com/item?id=16521385

It's not really an accessibility issue so much as a performance one. If it was a reasonable alternative to something like the rsync daemon I'd use it all the time. Unfortunately the performance issues and overhead is just too much.

IPFS isn't really an alternative to rsync, but the rest of your point stands.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#107
Maybe we should print this out on acid-free paper-thin flexible wood-pulp sheets stitched to together to form linear organized aggregations. Each aggregation would contain one or more works and be searchable using a SQL-like database. To make this plan really work there would need to be a collection of geographically distributed long term physical repositories that would receive periodic updates as new material became available.

All joking aside, I do wonder wither digital or analogue formats are better able to survive into the distant future.

* What impact will DRM have on the accessibility of our knowledge to future historians?

* Is anything recoverable from a harddrive or flash media after 500 years in a landfill?

* Will compressed files be more of less recoverable? What about git archives?

* Will the future know the shape of our plastic GI Joes toys but not the content of the GI Joes cartoon?

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#108
post #58

Earlier quoted context omitted.

Is there some tool or documentation describing this practice?

I'm sure someone could get you the info to get setup as a seeder. For modern clients it's rather rather trivial to manage that many torrents. Get any decent modern CPU, 4gb+ ram, and $560 in storage and you're off.

I think the problem is that because of the size of each torrent, and there's 1000 of them, it's difficult to effectively seed all at once, so instead people would rather seed sections at once, and rotate through them.

I'm not sure how people setup the rotation though, that can't be an incredibly common feature but I could be wrong.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#109

I don't see anyone having mentioned the possibility of posting this data to Usenet at all - at minimum for archival purposes which should be good for ~8-9 years. That way at least the data isn't lost. With so many of those torrents have 0 or 1 seed, this is a serious risk I think, despite the comments elsewhere about people rotating what they seed. I realize that doesn't solve the access problem for most people as mo…

Two thoughts on that. Encoding it to a text format with CRC data for posting to usenet is highly inefficient in terms of data storage. And 33TB of stuff is not going to be retained for 8-9 years, the last I checked due to the huge volume of binaries traffic, the major commercial usenet feed providers have at most 6-9 months of retention for the major binary groups. Beyond that it becomes cost prohibitive for them in terms of disk storage requirements. This is not an issue for the majority of their customers, 6-9 months is more than long enough retention to go find a 40GB 2160p copy of some recently-released-on-bluray movie.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#110
post #25

Related: Looking at harddisk cost per terabyte, quite often extern drives are cheaper than internal ones. For example right now in Germany I can get a WD 8TB USB 3.0 drive for 135€ but the cheapest internal 8TB drive costs 169€. Any idea why? It's puzzling.

For me on amazon.de the WD 8TB USB 3.0 drive is currently at EUR 159.99. Where do you get it for EUR 135?
Post reply on HN