Live data from Hacker News

How to circumvent Sci-Hub ISP block

fragile-credences.github.io

111–120 of 196 posts

Re: How to circumvent Sci-Hub ISP block

#111
post #16

What is forcing these UK ISPs to block these IP ranges?

Many years ago, Hollywood and the Music Industry argued in court in the UK that since most UK ISPs have a mechanism to block stuff like child pornography, they ought to be compelled by the court to use this same mechanism to protect the interests of these commercial entities.

The courts agreed. If your ISP has such a capability the court will cheerfully give rights holders authority to demand the ISP blocks stuff that they claim rights over.

The ISPs could choose to just not block stuff. I am looking at Sci-Hub right now, because my UK ISP (Andrews & Arnold) doesn't block anything. From time to time, parliamentarians get vexed about this, but there is a cultural memory in the place that passing laws intended to stop people from saying things doesn't work. Lady Chatterley's Lover was banned, because, the government argued, it was obscene but it turns out that a court didn't buy this argument, and instead Penguin made piles of money because everybody wanted to read the banned book.

But most UK ISPs have decided that it is in their better interests to block things. Their options for attempting to do this have narrowed over time, once upon a time DNS blocking was pretty effective, I assume that by now they're mostly relying on IP blocking, which of course means they run the risk of exciting collateral attacks...

Re: How to circumvent Sci-Hub ISP block

#112

Earlier quoted context omitted.

In Internet scale it's not a lot of data. Most people who think they have big data don't. Estimates I've seen put the total Scihub cache at 85 million articles totaling 77TB. That's a single 2U server with room to spare. The hardest part is indexing and search, but it's a pretty small search space by Internet standards.

The entire Library of Congress books collection is on the order of 40 million items. At 5 MB per book, this works out to about 200 TB of disk storage. At about $12/TB, hosting the entire LoC collection would cost roughly $2,400 presently, with prices halving about every three years.

Note that $2,400 is disks alone. You'd obviously need chassis, powere supplies, and racks. Though that's only 17 12 TB drives.

Factor in redundancy (I'd like to see a triple-redundant storage on any given site, though since sites are redundant across each other, this might be forgoable). Access time and high-demand are likely the big factor, though caching helps tremendously.

My point is that the budget is small and rapidly getting smaller. For one of the largest collections of written human knowledge.

There are some other considerations:

- If original typography and marginalia are significant, full-page scans are necessary. There's some presumption of that built into my 5 MB/book figure. I've yet to find a scanned book of > 200MB (the largest I've seen is a scan of Charles Lyell's geology text, from Archive.org, at north of 100 MB), and there are graphics-heavy documents which can run larger.

- Access bandwidth may be a concern.

- There's a larger set of books ever published, with Google's estimate circa 2014 being about 140 million books.

- There are ~300k "conventionally published" books in English annually, and about 1-2 million "nontraditional" (largely self-published), via Bowker, theh US issuer of ISBNs.

- LoC have data on other media types, and their own complete collection is in the realm of 140 million catalogued items (coinciding with Google's alternate estimate of total books, but unrelated). That includes unpublished manuscripts, maps, audio recordings, video, and other materials. The LoC website has an overview of holdings.

Published document scarcity is entirely imposed.

Re: How to circumvent Sci-Hub ISP block

#113
post #87

Earlier quoted context omitted.

> Cease to supply this system with the fruits of your research labors. Historically academics have felt forced to support this system, because for-profit journals are the high-prestige ones they must publish in in order to get tenure. This has changed for certain fields, but it isn’t as simple as just suggesting that one publish elsewhere.

What are some of the fields where this is changing?

For some fields the for-profit problem never happened at all. For example, in some branches of linguistics, history and archaeology the main journals have always been published by the same non-profit learned societies for decades (since the 19th century, sometimes). Prices for the hardcopy were always reasonable, and with the digital era, these journals became open access.

In other branches of those disciplines, I have seen that recently some big-name editors have founded new open-access journals with the express aim of gradually taking prestige away from for-profit journals. See here [0] (PDF).

[0] https://scholarsarchive.library.albany.edu/cgi/viewcontent.c...

Re: How to circumvent Sci-Hub ISP block

#114

By the way, Sci-Hub has stopped adding new articles to the database for a few months now (background: https://www.reddit.com/r/scihub/comments/mk46x4/scihub_v_els... ). It would be great to develop a truly decentralised solution. Having a database of individual torrent links for each paper might be a start.

Millions of individual torrents is not a great solution. Keeping them all seeded is basically impossible unless they run a seed for each one, at which point they might as well just host the files. Plus you'll never get the economy of scale that makes BitTorrent really shine. When you have a whole lot of tiny files that people will generally only want one or two of there isn't much better than a plain old website. A t…

Maybe Usenet? It already support massive copyright infringement yet it is still around.

Re: How to circumvent Sci-Hub ISP block

#115
post #23

Just setup a VPN on some cheap cloud provider. There are lots of sites UK ISPs block even though the sites themselves are not illegal or host illegal content. For e.g. torrent indexing services (the content itself may be illegal but purely providing a search across that content is basically doing what Google do). The UK internet is heavily filtered/censored and so doing this is useful anyway. Business ISP connections…

I travelled last week, and was horrified by how much is blocked by the mainstream ISPs in the UK. Afaik, my (London) ISP does not block anything. No idea why, as all the others quote high court orders.

Open Rights Group run https://www.blocked.org.uk/ to make blocking more visible and prompt action to contact ISPs about cases of overblocking.

You can enter any domain to test it, the data is collected by volunteers running an automatic tool.

Re: How to circumvent Sci-Hub ISP block

#116
post #57

Earlier quoted context omitted.

Yes, due to allocation of property rights. Cease to supply this system with the fruits of your research labors.

> Cease to supply this system with the fruits of your research labors. Historically academics have felt forced to support this system, because for-profit journals are the high-prestige ones they must publish in in order to get tenure. This has changed for certain fields, but it isn’t as simple as just suggesting that one publish elsewhere.

Stop supporting a system you won't inherit.

Re: How to circumvent Sci-Hub ISP block

#117

A few years back, frustrated with increasaed DNS blocking of Sci-Hub, I wrote a quick DNSMasq hack (haq?) to return Sci-Hub IPs for any "sci-hub. " possible. The shins-n-grits factor of surfing "scihub.elsevier.com" were palpable. https://old.reddit.com/r/Scholar/comments/7m3uin/meta_if_you... As others have mentioned, Sci-Hub also maintains a Tor presence, and you can access the Onion link using the Tor browser (pro…

Unfortunately the .onion is currently timing out.

Love the haq!

Re: How to circumvent Sci-Hub ISP block

#118
post #68
post #4

How do the alternative domains fare in the UK? With https://sci-hub.st/ I can circumvent the ISP block in Sweden successfully.

Works for me on Plusnet in the UK, thanks!

Interestingly it works on HTTPS but with HTTP you get a block page referencing the court order (confirmed by a friend on Plusnet)

See also: https://www.blocked.org.uk/site/http://sci-hub.st

Re: How to circumvent Sci-Hub ISP block

#119

By the way, Sci-Hub has stopped adding new articles to the database for a few months now (background: https://www.reddit.com/r/scihub/comments/mk46x4/scihub_v_els... ). It would be great to develop a truly decentralised solution. Having a database of individual torrent links for each paper might be a start.

from what I understand, the Authors are still free to send you their papers. So perhaps the simplest decentralized system is for each author to run an automatic email request system and have feeds / aggregators of papers titles/abstracts with author emails to make it easy to get the papers you want?

Re: How to circumvent Sci-Hub ISP block

#120
post #23

Earlier quoted context omitted.

I travelled last week, and was horrified by how much is blocked by the mainstream ISPs in the UK. Afaik, my (London) ISP does not block anything. No idea why, as all the others quote high court orders.

Many UK ISPs have "adult content" filters, which tend to be wide-reaching and block a lot more than just porn sites. But these are optional and can be turned off very easily. There's a smaller set of non-optional blocks (pirate/torrent sites) which you need a VPN to get around.

My ISP does not block any of the non-optional blocks either..

While I do use a VPN at times, it is never to get around blocks.

Although I don't use my ISPs dns, so if that is how they're attempting to comply I wouldn't notice.

Post reply on HN