Live data from Hacker News

Sci-Hub statistics and database

sci-hub.ru

101–110 of 150 posts

Re: Sci-Hub statistics and database

#101

Earlier quoted context omitted.

It is note worthy that most of physics (at least high energy physics) ist published on arxiv.org and open access. I don't know if sci hub bothers with publications that are available freely from an official source.

> published on arxiv.org and open access Don't use the term "open access" like this. A paper published on arXiv is free to read, and was freely published. "Open access" is a scam by the big publishers, where they don't take money from the readers , but make the authors pay. Or, putting it another way, anyone can pay their way in those journals and publish (sometimes sub par) papers.

Open Access has a precise definition:

https://openaccess.mpg.de/Berlin-Declaration

Publishers do misuse the concept though. They try to stay clear of using the term when they do. They use terms like Free Access or some of the more dubious colour variants of OA that doesn't provide all the freedoms that the Berlin Declaration of OA defines.

Interestingly articles uploaded to arXiv with the arXiv.org perpetual non-exclusive license are not OA as the reader is not allowed to redistribute the paper.

Re: Sci-Hub statistics and database

#102
post #98

Earlier quoted context omitted.

No “firm words” that I heard in my neck of the USA woods. I was struck by very low number of German downloads. Did I miss something?

No idea about Germany, but they may well be more law abiding?

No, that's absurd. More likely a different factor. Some googling and lo and behold: starting with Germany, the universities of several European countries cancelled their agreements with Elsevier, pressuring Elsevier to give them better deals, including open access by default. I imagine it lessens the need for Sci-Hub.

Re: Sci-Hub statistics and database

#103
post #26

It's interesting how sci-hub's papers on medicine dwarf those in many other fields like comp-sci, math, and physics. I wonder if that reflects the number of papers in those fields, or if sci-hub just has a non-representative sample. If the latter, why?

I guess today is much easier to find new noteworthy, publishable facts in medicine than physics. New diseases are discovered every year, and old diseases are poorly understood (e.g Alzheimer disease), and the treatments for many of them are still sub-optimal, or even inexistent. Every patient is different, individual cases are research-worth. We only got antibiotics in the 1940s. On the other hand, most big breakthro…

> much easier to find new noteworthy, publishable facts in medicine than physics

Plus I would imagine that everyone wants their illness looked into, thus that's where funding tends to go. I care much less for physicists to figure out what dark matter is than how to treat health problem X that bothers me daily. (Just an example. In my particular case I'm healthy and would actually be quite excited about dark matter findings compared to any individual illness solution... but still.)

This comment would be worth a lot more with some stats about funding going towards the different fields, though. Not sure where to find that.

Re: Sci-Hub statistics and database

#104

Anyone notice the logo update? The key loop has been exchanged for a hammer and sickle.

I did not, that's somewhat interesting (not hard to explain if you read the about pages, though). It's also well within the domain of politics and that seems to be something HN generally avoids.

Re: Sci-Hub statistics and database

#105
post #89

The publishers now encode their papers with individual identifiers, that generationP calls UUIDs on all pdf's. That means they can trace it back to the institution(and perhaps the actual prof?) They have then sent nastygrams to threaten them with fees or loss of access to the institution - potentially serious punishment. What is needed is a way to run the papers through an OCR recognition program to create renewed te…

Makes me wonder how much they'd have to offer me to accept the task of implementing such tracking algorithms to modify other people's scientific papers for the purpose. Certainly I won't be the cheapest, but still. Either some developer is very vested in the idea of keeping science a secret or someone got a very nice bonus.

Edit: also the sysadmin that keeps this database safe without 'accidental' data loss on UUID to downloader mappings.

Re: Sci-Hub statistics and database

#106
post #89

The publishers now encode their papers with individual identifiers, that generationP calls UUIDs on all pdf's. That means they can trace it back to the institution(and perhaps the actual prof?) They have then sent nastygrams to threaten them with fees or loss of access to the institution - potentially serious punishment. What is needed is a way to run the papers through an OCR recognition program to create renewed te…

> I am not sure how Sci-hub can get past this

Quite trivially, actually, thanks to the good old analogue hole: https://news.ycombinator.com/item?id=30084193

All sci-hub would have to do in this case is download the same paper through three or more accounts (different institutions, networks, countries?) at three different times, rasterize them and keep the common denominator. If a pixel has no common denominator, they'd have to fall back to a default value. This is by no means a perfect method and it has its weaknesses and pitfalls, but the resulting PDF will be far less useful as a means to de-anonymize accounts using information from the PDF itself.

Publishers still have other sources of information to de-anonymize accounts if the multiple accounts/downloads aren't truly isolated from one another.

Re: Sci-Hub statistics and database

#107

Alexandra Elbakyan is a titan and a saint. I couldn't have been able to finish my research without access to papers my institution wasn't subscribed to.

Isn't it fantastic that we are alive and seeing the resurrection of the great library of Alexandria right before our eyes? She has done more than any other organization or individual in the history of mankind when helping people in second and third world country pursue advance research since the advent of internet. Well she and the people who pirate and distribute MS Office. Faculties around the world recommend scihu…

The great library of Alexandra, you mean.

Re: Sci-Hub statistics and database

#108
post #89

The publishers now encode their papers with individual identifiers, that generationP calls UUIDs on all pdf's. That means they can trace it back to the institution(and perhaps the actual prof?) They have then sent nastygrams to threaten them with fees or loss of access to the institution - potentially serious punishment. What is needed is a way to run the papers through an OCR recognition program to create renewed te…

Do a reflow of the entire document? As long as it's semantically recognizable it can be done. Figures are harder. ML can really help here: someone make a pdf2tex!

Re: Sci-Hub statistics and database

#109

It's interesting how sci-hub's papers on medicine dwarf those in many other fields like comp-sci, math, and physics. I wonder if that reflects the number of papers in those fields, or if sci-hub just has a non-representative sample. If the latter, why?

[deleted]
Post reply on HN