Live data from Hacker News

Archivists Are Trying to Make Sure LibGen Never Goes Down

vice.com

21–30 of 270 posts

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#21
post #3

There is a huge amount of duplication there (i.e. books that have many scans), I wonder if it would be better to tackle that versus doing a straight backup.

I think the duplication issue is probably overstated. I doubt tackling that would shave off more than 20% of the total backup size.

It's probably more of a nuisance for people wanting to use the content. E.g., copies with different metadata or tags.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#22
post #5

one of the next interplanetary or Interstellar Probe should carry a copy of the sci-hub torrent in some kind of permanent storage

Storing that amount of information in a way that an unknown alien species would be able to read (even assuming technical expertise greater than our own) is a huge problem.

Keep in mind that they don't know our written or computational language and there's nothing about our technology that is inherently self-explaining/obvious.

Even the assumption that they'd use binary computers (rather than trinary, or other technology not based around electrical voltages) is open to debate.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#23
post #11
post #7

Is there a way to just download the whole 32TB to your own machine? I see a ton of mirrors but the content seems to be highly fragmented between them

There are ways to do so. The archive is made up of many, many torrents (I believe it's a monthly if not biweekly update of the database). If you have the storage/bandwidth availability for the whole 32TBs, please get in touch and I may be able to help you get the whole deal without too much hassle. Otherwise, just pick some torrents (it would be best to pick them based on torrent health, but they are so many to check…

I'm pretty surprised by the lack of seeders. Out of the 2438 torrents listed, a third have 0 seeders, another third have 1 seeder, and all but 5 have less than 10. Hopefully the publicity boosts those numbers.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#24
post #2

This is an extremely important effort. The LibGen archive contains around 32 TBs of books (by far the most common being scientific books and textbooks, with a healthy dose of non-STEM). The SciMag archive, backing up Sci-Hub, clocks in at around 67 TBs [0]. This is invaluable data that should not be lost. If you want to contribute, here's a few ways to do so. If you wish to donate bandwidth or storage, I personally k…

> Lastly, you can always contribute books. If you buy a textbook or book, consider uploading it (and scanning it, should it be a physical book) in case it isn't already present in the database.

There's no easy solution for scanning physical books, is there?

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#26
post #19

Why not publish the site over IPFS, that would make P2P hosting much simpler?

Currently (at least for the-eye) it's about IPFS's barrier of entry. I expect LibGen's case to be similar. Most people don't know about it, and if even those that knew about it had to learn how IPFS works etc, they would probably just try to find the book they're looking for elsewhere.

True, I too find it not ideal, but having such a massive library available over it surely would increase the interest in lowering the barrier of entry?

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#27
post #2

This is an extremely important effort. The LibGen archive contains around 32 TBs of books (by far the most common being scientific books and textbooks, with a healthy dose of non-STEM). The SciMag archive, backing up Sci-Hub, clocks in at around 67 TBs [0]. This is invaluable data that should not be lost. If you want to contribute, here's a few ways to do so. If you wish to donate bandwidth or storage, I personally k…

> Lastly, you can always contribute books. If you buy a textbook or book, consider uploading it (and scanning it, should it be a physical book) in case it isn't already present in the database. There's no easy solution for scanning physical books, is there?

There are providers [1] that will destructively scan the book for you and return a PDF. If you want to preserve the book, you're stuck using a scanning rig [2]. The Internet Archive will also non-destructively scan as part of Open Library [3], but they only permit one checkout at a time of scanned works, and the latency can be high between sending them a book and it becoming available. FYI, 600 DPI is preferred for archival purposes.

[1] http://1dollarscan.com/ (no affiliation, just a satisfied customer, can't scan certain textbooks due to publisher threats of litigation)

[2] https://www.diybookscanner.org/

[3] https://openlibrary.org/help/faq

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#28
post #2

This is an extremely important effort. The LibGen archive contains around 32 TBs of books (by far the most common being scientific books and textbooks, with a healthy dose of non-STEM). The SciMag archive, backing up Sci-Hub, clocks in at around 67 TBs [0]. This is invaluable data that should not be lost. If you want to contribute, here's a few ways to do so. If you wish to donate bandwidth or storage, I personally k…

Is it an important effort, though? If the LibGen archive disappeared, I doubt my life would change in any meaningful way.

I'd love to be proven wrong.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#29
post #2

This is an extremely important effort. The LibGen archive contains around 32 TBs of books (by far the most common being scientific books and textbooks, with a healthy dose of non-STEM). The SciMag archive, backing up Sci-Hub, clocks in at around 67 TBs [0]. This is invaluable data that should not be lost. If you want to contribute, here's a few ways to do so. If you wish to donate bandwidth or storage, I personally k…

> Lastly, you can always contribute books. If you buy a textbook or book, consider uploading it (and scanning it, should it be a physical book) in case it isn't already present in the database. There's no easy solution for scanning physical books, is there?

[deleted]

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#30
post #19

Why not publish the site over IPFS, that would make P2P hosting much simpler?

Currently (at least for the-eye) it's about IPFS's barrier of entry. I expect LibGen's case to be similar. Most people don't know about it, and if even those that knew about it had to learn how IPFS works etc, they would probably just try to find the book they're looking for elsewhere.

I am not fully aware how IPFS operates, but wouldn't it at least solve the back-end mirroring? Front-end servers would then "only" need to access IPFS for continuous syncing of metadata (for search) and fetching user-requested files (upon request).
Post reply on HN