Live data from Hacker News

Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

pilimi.org

71–80 of 438 posts

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#71

Earlier quoted context omitted.

Do you really want your ISP to know which piracy sites you frequent? This is all being sent in plain text. Or they could change the content, insert a redirect, or inject ads without your knowledge. TLS is needed on all websites - not just those with interaction.

https won't keep your ISP from knowing you visited the site. And the rest of those? For a text-only blog, they seem kinda trivial.

Not disagreeing but presenting a hypothetical:

If the user requests the page from Internet Archive, Common Crawl or even Google Cache, how does the ISP know what the user requested. (NB. Neither IA nor Google Cache require sending SNI,^1 so the ISP may only see IP addresses).

With IA, the IP address alone does not reveal which IA site or page the user is requesting. There is more to IA than only Wayback Machine.

With Common Crawl, the user can send the Cloudfront domain name instead of a commoncrawl.org domain. Are all ISPs going to know that this is Common Crawl. Even if they expend the effort to learn, what benefit is achieved.

With Google Cache, the IP address alone does not reveal which Google site the user is accessing. Needless to say, there are many, many domains using these IP addresses.

There is nothing that requires any web user to retrieve web pages from a given host. The page may be mirrored at a number of hosts. Some of those hosts might offer HTTPS, support TLS1.3 and not require plaintext SNI/offer encrypted ClientHello.

Even assuming an ISP can determine what domain name a customer is sending in a Host header or ClientHello packet, it would still be necessary to subpoena the archive/CDN/cache to figure out precisely what pages were being requested.

1. The same party is controlling all the server certificates. IA controls the certificates for all IA domains, Amazon (issues and) controls all the certificates for Cloudfront customers and Google controls all the certificates for Google domains. Perhaps there are web users commenting on HN who believe that ingress/egress traffic for site saved/hosted/cached at an archive/CDN/cache is somehow private as against the company running the archive/CDN/cache in a meaningful way. I am not one of them.

As for the question of an ISP modifying the contents of web pages, this is an issue that could be addressed contractually in a subscriber agreement. It stands to reason that if this was a serious issue and not merely a hypothetical one raised by nerds debating the merits of TLS then it would be addressed in such agreements.

As for the "injection of advertising" issue as a argument in favour of the way TLS^2 is being administered on the web, IMO this is a bit silly since (a) it is trivial to filter such advertising (e.g., Javascript in the examples I saw) out out of the page and/or block it from running/connecting/loading and (b) the amount of "tech" company-mediated advertising that web users endure in spite of using TLS is enormous. More likely than being seen as a threat to web users, the injection of advertising by ISPs was seen as a threat to the advertising revenue of "tech" companies. The later are responsible for facilitating the injection of advertising (by their customers, not their competitors, i.e., ISPs), not preventing it.

2. By "TLS administration" I do not mean encryption as a concept nor certificates as a concept. I mean TLS administration measures designed to support "tech" companies first and web users second, if at all. A system where the questions of "threat model" and "trust" are both decided by "tech" companies not users.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#72

Earlier quoted context omitted.

The homepage has a link to the onion site.

So why does it have a clearnet address? To have more reach? What’s their threat model such that a clearnet presence could possibly out the people behind this?

I don't know. I'd guess reach. Probably wouldn't be on HN if it were darknet only.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#73

I wonder if there are any search engines dedicated to indexing these kinds of libraries. I know there's a decent one just for scihub, but it would be awesome if I could do a Google-style search that returned the contents of books, magazines and journal articles instead of just websites.

Book metadata is widely available via sites like e.g. Open Library. With good metadata, full text search is not as relevant.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#74
post #30

HTTP-only makes me weary of visiting a self-professed piracy site. They couldn't even spring for a Let's Encrypt cert?

*wary For some reason I'm seeing this mistake more and more lately. https://en.wiktionary.org/wiki/weary vs https://en.wiktionary.org/wiki/wary

People have mixed up leery and wary and it has memetically become ‘weary’.

I notice it more and more too.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#75

To be fair, Z-Library doesn't charge unless you want to download more than 10 books per 24 hour period. That's per account and although they ask you not to open multiple accounts they don't seem to do anything to stop you.

What kind of fairness can there be in charging for stolen books? I believe in free access to education, but charging for these books they have no rights to is a whole other thing.

Technically, they're not charging for the books, they're charging for the bandwidth.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#76

It's really funny to think about how the advances of technology keeps changing how we perceive books. 7TB is even a commodity disk these days. And it's a lot less than the torrent of scientific papers that floated around some time ago (that was ~18TB IIRC).

I foresee storage density reaching the point that for most ordinary people "online" becomes rather unimportant. What would be the effects of technology when computers behave as in early science fiction, as stand-alone oracles? [1] [1] https://www.timeshighereducation.com/opinion/2048-informatio...

How would the appeal of streamers and live data/content settle out in that case? Sometimes context is available in the moment that makes it easier for all parties to consume and analyze in that moment as well.

Since transient, ethereal meme culture is also basically emergent culture now, it's difficult not to also foresee a greater cultural divide in such a case. This is saying nothing of live data tools as well, even weather data...

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#77

To be fair, Z-Library doesn't charge unless you want to download more than 10 books per 24 hour period. That's per account and although they ask you not to open multiple accounts they don't seem to do anything to stop you.

What kind of fairness can there be in charging for stolen books? I believe in free access to education, but charging for these books they have no rights to is a whole other thing.

I don't quite agree. I mean, they provide useful service, and it costs money to run it. It's ok that they earn (even if it's actually making a profit, not just covering the costs).

That being said, 10 downloads/day feels a bit restrictive to me. I'd get if it was 100, or 50, heck, maybe even 20. I mean, I don't appreciate that it's not mirrorable in the first place, but maybe they cannot afford it, I don't know... But 10 feels less than somebody researching a new topic might need to access in a day, even if he won't read them all immediately.

…That being said as well, it has some really nice UI. I wish somebody did it for Libgen.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#78
post #61

Earlier quoted context omitted.

Wasn't that what google books was supposed to be?

I assume Google abandoned this along with all their earlier mission statements in favour of building another chat app.

I can't imagine a life at google. So much promise, all turned to ash.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#79
post #30

HTTP-only makes me weary of visiting a self-professed piracy site. They couldn't even spring for a Let's Encrypt cert?

*wary For some reason I'm seeing this mistake more and more lately. https://en.wiktionary.org/wiki/weary vs https://en.wiktionary.org/wiki/wary

I've heard people say it. People who I think should know better. It's really frustrating.

Re: Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)

#80
post #61

Earlier quoted context omitted.

Wasn't that what google books was supposed to be?

I assume Google abandoned this along with all their earlier mission statements in favour of building another chat app.

Could you take a moment to check it Google Books search still exists?

I'll give you a hint: https://books.google.com/?hl=en

> Search the world's most comprehensive index of full-text books.

Post reply on HN