Live data from Hacker News

Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

pilimi.org

101–110 of 438 posts

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#101
post #77

Earlier quoted context omitted.

What kind of fairness can there be in charging for stolen books? I believe in free access to education, but charging for these books they have no rights to is a whole other thing.

I don't quite agree. I mean, they provide useful service, and it costs money to run it. It's ok that they earn (even if it's actually making a profit, not just covering the costs). That being said, 10 downloads/day feels a bit restrictive to me. I'd get if it was 100, or 50, heck, maybe even 20. I mean, I don't appreciate that it's not mirrorable in the first place, but maybe they cannot afford it, I don't know... Bu…

It's a good thing their hosting provider is okay with providing bandwidth per individual user account on the site of up to—how many books did you say again, 50?

Without sarcasm, I don't think the bandwidth bill cares that you find it restrictive. That even more than a handful are free every day for every account is honestly a lot, since that means virtually nobody will need to contribute to the costs they're collectively incurring. And if you're unable to pay, you can still skim a few dozen books (making two or three accounts isn't that hard to do by hand) every day, and go back to any you've already downloaded previously too. And offer them to friends to offload the server.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#102
post #89

Earlier quoted context omitted.

What kind of fairness can there be in charging for stolen books? I believe in free access to education, but charging for these books they have no rights to is a whole other thing.

At this point is worth noting that there are reputable sources for free books, such as public libraries and Project Gutenberg. I realize that neither source will satisfy many of the people on HN, simply because there is a need for current technical books.

What on earth does "reputable" mean here?

Total compliance with the regulatory capture of publishing companies? Full support of the landgrab claims of the Disney corporation et al?

You don't have to be an anarchist to look at the status quo and think some amount of civil disobedience is the correct, proper and right thing to do. Also believing that a claim, fully Disney supported, that such an amount of civil disobedience is somehow ethically evil is, in fact, somewhat disreputable. As disreputable as the insanely high journal subscription fees for taxpayer funded research, for example.

The publishing companies chose this path willingly and with prejudice for their profit turning the relevant law against the people. Are they reputable given they did so? It's hardly an outlying position around here to think they really aren't anything of the sort. Refusing to accept that on mass, until appropriate reform is supported and enacted could be considered quite worthwhile.

You may of course, disagree.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#103

Earlier quoted context omitted.

> they can serve you anything they want. Great. More books! No really, I don't understand this argument. A static site served by plain http is perfectly appropriate. It's like a poster hanging on the wall for all to see. Of course people can paint over it, but it doesn't really matter.

They could serve you javascript that exploits your browser. At the very least, they could replace that bitcoin donation address with their own. That's a tempting target if nothing else.

If you think you're high value enough to have someone target you specifically by getting on your LAN or gaining access to (or coercing) an upstream ISP to serve you a browser 0-day reachable only by laying in wait for you to visit an HTTP site because there is no other way in, that's not going to be for a free books website.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#104
post #27

So, if I get it right. First there was Libgen, which is mirrorable. Then, some Z-Library copied Libgen and added some more books, without making it mirrorable. The goal is to make these new books, which are not mirrorable — mirrorable (i.e. to "preserve" them). So, why not just re-upload them to Libgen, then? I guess somebody will do that now anyway, but you could easily done it in the first place, without making you…

From their FAQ: > Q: Should the Z-Library collection be added to Library Genesis? > A: Yes! However, it is tricky. Library Genesis splits out its collection between non-fiction and fiction. They also have relatively high quality standards. If you are interested in organizing all the books to meet their requirements, let us know.

Can't they just tag these books as "not reviewed yet"? Anyone looking for these books can then just decide to include them.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#105

Earlier quoted context omitted.

> they can serve you anything they want. Great. More books! No really, I don't understand this argument. A static site served by plain http is perfectly appropriate. It's like a poster hanging on the wall for all to see. Of course people can paint over it, but it doesn't really matter.

HTTP connections can be used as a weapon against others. One example is China’s Great Cannon. https://citizenlab.ca/2015/04/chinas-great-cannon/

That's quite dated by now. If you are in a position to inject traffic, you are likely also able to simply use that uplink to send traffic of your own. I'd be surprised if this is still in use, especially outside of China (or a poor not-so-tech-savvy country like North Korea) and wasn't just a quick hack at the time.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#107

Earlier quoted context omitted.

> https won't keep your ISP from knowing you visited the site If you use DoH, yes it does. Unless I'm mistaken. They only know the IP address of the remote server.

And nobody would ever think of keeping a reverse DNS index.

With shared hosting, NAT, etc. there may be several sites sharing the same IP.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#108
post #99

Earlier quoted context omitted.

It was paralyzed by legal disputes with book publishers. In the years the lawsuits were going on, nearly everyone left the project. And then the lawyers have put in so many red lines that it's nearly impossible to make any changes to it.

Yup, I read about that on Wikipedia, but I can't help but not care. If a company touts massive initiatives and then gets bogged down in lawsuits, it seems like they didn't do the basic due diligence to avoid that. (Uber, AirBNB, and others seem to also have these headwinds, though not to the extent that it led to permanent paralysis, so maybe Google made a bet they thought they'd win and then didn't, whereas these ot…

This well-written article, "Torching the Modern-Day Library of Alexandria"[0], might change your mind on that. I found it to be a compelling and tragic story.

Edit: actually, I think this would make a good submission. Looks like it hasn't been posted since 2017.

[0]: https://www.theatlantic.com/technology/archive/2017/04/the-t...

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#109
post #11

7TB of compressed text? I don't think humanity has generated that much written words in it's entire existence. Although it would be an interesting Fermi Problem to estimate (and don't forget just how well text compresses). This has to be a lot of duplicates or bad formats (images). This would be far more useful to people with some curating.

(Ignoring, as someone else already pointed out, that this is PDFs with images and scans probably, not just plain text... I got curious and did some napkin math about your claim)

The top ten countries average north of a hundred thousand books per year each[1], so let's say a million books per year globally because I haven't got the time to extract the numbers and sum them all up exactly. It probably wasn't as much fifty years ago, but we're also completely ignoring the Internet which is way bigger because of user-contributed content (not to mention things like newspapers, meeting notes, etc.), so let's say this was the case for the past fifty years. That's fifty million books. Average book has 85k words[2], and a word is like 5 characters on average. Not all countries write in English, especially some big ones like India and China, so we could probably double it on a global scale but let's go for a conservative 7.5 bytes per word and add a byte for the space (I'm ignoring other punctuation). That comes out to 8.5×85e3×1e6 which is less than a gigabyte and uncompressed. Decent compression iirc makes it a fifth of the size, so 145MB compressed.

I've still got to be an order of magnitude off because we write a whole lot more than just books (and even if we were just looking at books: there are also book revisions, drafts, etc. I'm just counting the published words), but that's still less than expected.

In conclusion, it comes out to about 1/50'000th of 7TB compressed, which I would say lends some credence to the claim—even if I feel like I must have made a mistake somewhere because 145MB for ~all books of the past 50 years from all countries seems quite little.

[1] https://en.wikipedia.org/wiki/Books_published_per_country_pe...

[2] a few sources on ddg, e.g. https://www.tckpublishing.com/how-long-should-a-book-be/

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#110
post #92

Earlier quoted context omitted.

So why does it have a clearnet address? To have more reach? What’s their threat model such that a clearnet presence could possibly out the people behind this?

There is no ssl so no cert fingerprint shodan matching leak

Mind elaborating?
Post reply on HN