Live data from Hacker News

Archivists Are Trying to Make Sure LibGen Never Goes Down

vice.com

61–70 of 270 posts

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#61
post #2

This is an extremely important effort. The LibGen archive contains around 32 TBs of books (by far the most common being scientific books and textbooks, with a healthy dose of non-STEM). The SciMag archive, backing up Sci-Hub, clocks in at around 67 TBs [0]. This is invaluable data that should not be lost. If you want to contribute, here's a few ways to do so. If you wish to donate bandwidth or storage, I personally k…

[deleted]

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#62
post #30
post #19

Earlier quoted context omitted.

Currently (at least for the-eye) it's about IPFS's barrier of entry. I expect LibGen's case to be similar. Most people don't know about it, and if even those that knew about it had to learn how IPFS works etc, they would probably just try to find the book they're looking for elsewhere.

I am not fully aware how IPFS operates, but wouldn't it at least solve the back-end mirroring? Front-end servers would then "only" need to access IPFS for continuous syncing of metadata (for search) and fetching user-requested files (upon request).

[deleted]

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#63
post #33

Earlier quoted context omitted.

> an unknown alien species Not necessarily for aliens ... but why not keep a backup in a safe place outside the dangers of earthlings ? OTOH, I think a sufficiently advanced alien intelligence will be able to decipher the information structures we use regardless of differences in technology. It's possible though there will be missing links in that archive, which will need to be supplemented with a primary secondary a…

What about a satellite backup in orbit around earth? Maybe an elliptical orbit, coming around a few times a year or something

What format? I'm eager to know the solution to this non-problem.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#64
post #11
post #7

Is there a way to just download the whole 32TB to your own machine? I see a ton of mirrors but the content seems to be highly fragmented between them

There are ways to do so. The archive is made up of many, many torrents (I believe it's a monthly if not biweekly update of the database). If you have the storage/bandwidth availability for the whole 32TBs, please get in touch and I may be able to help you get the whole deal without too much hassle. Otherwise, just pick some torrents (it would be best to pick them based on torrent health, but they are so many to check…

If LibGen can announce all of the torrents in a JSON payload with health metadata, that can be consumed for automated seedbox consumption and prioritization. Check out ArchiveTeam's Warrior JSON project payload [1] for inspiration. It need not even be generated on-demand; render it on a schedule and distribute at known endpoints.

[1] https://warriorhq.archiveteam.org/projects.json

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#67
post #2

This is an extremely important effort. The LibGen archive contains around 32 TBs of books (by far the most common being scientific books and textbooks, with a healthy dose of non-STEM). The SciMag archive, backing up Sci-Hub, clocks in at around 67 TBs [0]. This is invaluable data that should not be lost. If you want to contribute, here's a few ways to do so. If you wish to donate bandwidth or storage, I personally k…

> Lastly, you can always contribute books. If you buy a textbook or book, consider uploading it (and scanning it, should it be a physical book) in case it isn't already present in the database. There's no easy solution for scanning physical books, is there?

There are DIY book scanners (http://diybookscanner.org) and products such as the Fujitsu ScanSnap SV600. The SV600 has decent features like page-detection and finger-removal (I recommend using a pencil's eraser tip). I have personally used it to scan dozens of books, with satisfactory results.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#68
post #56

Earlier quoted context omitted.

I just read the article and your comments here and I'm a bit unsure what's the difference to the Internet Archive. Is it that the IA can archive them but not make them public for legal reasons and The-Eye is more focused on keeping them online and accessible no matter what?

Yes. It is extremely likely IA has the LibGen corpus archived, but darked (inaccessible), to prevent litigation.

There are quite a few such copies, on the 'just in case' principle.

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#69
post #2

This is an extremely important effort. The LibGen archive contains around 32 TBs of books (by far the most common being scientific books and textbooks, with a healthy dose of non-STEM). The SciMag archive, backing up Sci-Hub, clocks in at around 67 TBs [0]. This is invaluable data that should not be lost. If you want to contribute, here's a few ways to do so. If you wish to donate bandwidth or storage, I personally k…

For important archives like this maybe we need some sort of turn-key solution for the masses? Like a Raspberry Pi image that maintains a partial mirror. Imagine if one could by a RPi and external HD, burn the image, and connect it to some random wifi network (at home, at work, at the library, etc).

Re: Archivists Are Trying to Make Sure LibGen Never Goes Down

#70

Why not publish the site over IPFS, that would make P2P hosting much simpler?

In my experience ipfs doesn't actually work. I'd love to be proven wrong, but the reason why nobody uses ipfs even when it seems like a great fit is bect it's not really usable.

This is my experience as well. In theory, IPFS is exactly the right thing for LibGen, but in practice I consider it unusable.
Post reply on HN