Live data from Hacker News

Archivists Are Trying to Make Sure LibGen Never Goes Down

vice.com

11–20 of 270 posts

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#11
post #7

Is there a way to just download the whole 32TB to your own machine? I see a ton of mirrors but the content seems to be highly fragmented between them

There are ways to do so. The archive is made up of many, many torrents (I believe it's a monthly if not biweekly update of the database). If you have the storage/bandwidth availability for the whole 32TBs, please get in touch and I may be able to help you get the whole deal without too much hassle. Otherwise, just pick some torrents (it would be best to pick them based on torrent health, but they are so many to check manually) and try to keep seeding as much as possible.

EDIT: To find libgen's torrents health, check out this google sheet: https://docs.google.com/spreadsheets/d/1hqT7dVe8u09eatT93V2x...

Thanks frgtpsswrdlame for the heads up.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#12
post #11
post #7

Is there a way to just download the whole 32TB to your own machine? I see a ton of mirrors but the content seems to be highly fragmented between them

There are ways to do so. The archive is made up of many, many torrents (I believe it's a monthly if not biweekly update of the database). If you have the storage/bandwidth availability for the whole 32TBs, please get in touch and I may be able to help you get the whole deal without too much hassle. Otherwise, just pick some torrents (it would be best to pick them based on torrent health, but they are so many to check…

Thanks! I don't have 32TB free locally at the moment but I might soon. If and when that happens, I'll get in touch :)

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#13
post #11
post #7

Is there a way to just download the whole 32TB to your own machine? I see a ton of mirrors but the content seems to be highly fragmented between them

There are ways to do so. The archive is made up of many, many torrents (I believe it's a monthly if not biweekly update of the database). If you have the storage/bandwidth availability for the whole 32TBs, please get in touch and I may be able to help you get the whole deal without too much hassle. Otherwise, just pick some torrents (it would be best to pick them based on torrent health, but they are so many to check…

Actually there is now a google sheet which shows the health of the torrents so it should be easy to pick the most helpful torrents. It's linked in this post: reddit.com/e3yl23

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#14
post #3

There is a huge amount of duplication there (i.e. books that have many scans), I wonder if it would be better to tackle that versus doing a straight backup.

I think the duplication issue is probably overstated. I doubt tackling that would shave off more than 20% of the total backup size.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#15
post #5

one of the next interplanetary or Interstellar Probe should carry a copy of the sci-hub torrent in some kind of permanent storage

Do we have anything rated for a few millennia of interstellar radiation besides etched gold plates?

Glass: https://news.microsoft.com/innovation-stories/ignite-project..., https://en.wikipedia.org/wiki/5D_optical_data_storage

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#17
post #5

one of the next interplanetary or Interstellar Probe should carry a copy of the sci-hub torrent in some kind of permanent storage

Do we have anything rated for a few millennia of interstellar radiation besides etched gold plates?

Microsoft's project Silica [0] may hopefully provide really long term, large capacity archive grade storage on earth. I wonder what effects interstellar radiation has on them.

[0] https://www.theverge.com/2019/11/4/20942040/microsoft-projec...

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#18

Earlier quoted context omitted.

This is a downside of Libgen: duplicate uploads, missing or erroneous metadata. You start wishing that there was at least some curation of the collection, so it could approach the quality of an academic library catalogue as many users are usedto. But I guess the people behind Libgen want to keep the number of people with database edit rights small. (When you upload a book, you yourself can edit the metadata for that…

Maybe they should consider a system where users can suggest tags/metadata or flag erroneous data that can be reviewed and allowed by a select few?

Integration with BookBrainz would be nice. The Brainz projects already consist of massive amounts of metadata curation and it would be possible to transfer that knowledge a bit.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#19

Why not publish the site over IPFS, that would make P2P hosting much simpler?

Currently (at least for the-eye) it's about IPFS's barrier of entry. I expect LibGen's case to be similar. Most people don't know about it, and if even those that knew about it had to learn how IPFS works etc, they would probably just try to find the book they're looking for elsewhere.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#20
post #3

There is a huge amount of duplication there (i.e. books that have many scans), I wonder if it would be better to tackle that versus doing a straight backup.

I think the duplication issue is probably overstated. I doubt tackling that would shave off more than 20% of the total backup size.

Speaking from personal experience, I usually see several results for any search. Granted, there's a big selection bias there, but 20% seems way too small.
Post reply on HN