I really hope sci-hub survives this. Sci-hub and libgen are like an entirely different internet, one allowing you to dive as deep as you wish into any technical subject. There’s really no comparison I have found anywhere for the depth of material available. People always point to Wikipedia, but all of that is surface level. If you to build something, research something, or just really delve into it, there’s no substi…
Yes, it’s a new and better world. On the other side of the $275 per-paper Elsevier paywall, there are researchers who wish more people would read their papers. In my experience digging into robotics kinematics, authors are happy to answer questions and can point me to the right person when I want to send a check to support investigating specific research questions. The paywall deceives; science is neither an institut…
Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate
151–160 of 171 posts
Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate
#152Earlier quoted context omitted.
I have been pleased that The Journal of Field Robotics is well represented on scihub. I have an open source off road robot I am designing and the journal is literally about robots out in fields and stuff. I am a "serious hobbyist" in that I believe my open source contributions to be at least somewhat helpful to others, but it's not the kind of thing that would justify paying for paywalled papers. I just want to glanc…
Wow! The content in JFR is fantastic! Thank you for the pointer!
Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate
#153I really hope sci-hub survives this. Sci-hub and libgen are like an entirely different internet, one allowing you to dive as deep as you wish into any technical subject. There’s really no comparison I have found anywhere for the depth of material available. People always point to Wikipedia, but all of that is surface level. If you to build something, research something, or just really delve into it, there’s no substi…
I really wish someone could build a better UI for this research internet. Hyperlinks for all references would be a good start. Finding some way to make some automatic glossary of definitions of technical terms would make scientific papers substantially more accessible too.
Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate
#154Torrent seeding effort: https://www.reddit.com/r/DataHoarder/comments/nc27fv/rescue_... All papers on sci-hub are available as torrents from library genesis. The full collection contains 85 million articles (before this announcement), and is about 80TB. If anything ever happens to sci-hub or library genesis, there's enough people out there with backups that a replacement can be set up fairly quickly, albeit without t…
Is it compressible or already compressed?
Recent publications are virtually always based on direct PDF renders, and tend to be a few 100 kB per article.
Older publications are often scanned from paper-based copies, and can be about 10-20x larger, depending on the source. These may or may not have OCRed text, and OCR itself may be of variable quality. For documents with images or diagrams, those also add to both size and difficulty in vectorising copies.
It's possible to go through larger scans and regenerate them as rendered PDFs. That's intensive and error prone. There's also a range of viewpoints on archival as to whether it's preferable to retain the full expression of the original published version (and often accumulated marginalia and other marks of a specific instance), or to optimise for both storage and automated processing through reprocessed renders. The costs are high (typically you'll require a human or multiple humans to proof each work), though the storage and line-transmission savings are considerable.
I lean toward the latter myself. The attitude of other archivists (notably the Internet Archive) is to capturing as faithful a replication of originally-published formats as possible, at considerable cost in both storage and accessibility. (This applies to the Archives work in print, online / Web, and other document formats.)
Pressed, I'd strongly recommend a "capture what you can, reprocess according to need and demand as possible" approach.
Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate
#155Earlier quoted context omitted.
No one cares. This thread is about free access to papers and not another paid service that forces you to pay monthly fees for something that could be a free service. In that sense you aren't any better than large online publishers. 8 bucks a month for a scientific paper search engine? Really?
Hiya, Well, I definitely agree with your sentiment in a normative sense that scientific papers should be free and readily accessible to all -- in part because a lot of it is funded through tax dollars! But given the current state of affairs, we're looking at making that information accessible to people without having to pay exorbitant fees to access individual research. We also offer steep discounts for students or a…
Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate
#156Torrent seeding effort: https://www.reddit.com/r/DataHoarder/comments/nc27fv/rescue_... All papers on sci-hub are available as torrents from library genesis. The full collection contains 85 million articles (before this announcement), and is about 80TB. If anything ever happens to sci-hub or library genesis, there's enough people out there with backups that a replacement can be set up fairly quickly, albeit without t…
Is there any legal risk to users in North America that do this? Is this copyrighted material?
2. seed from your seedbox, not your personal device
3. profit
Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate
#157Earlier quoted context omitted.
Yes, it’s a new and better world. On the other side of the $275 per-paper Elsevier paywall, there are researchers who wish more people would read their papers. In my experience digging into robotics kinematics, authors are happy to answer questions and can point me to the right person when I want to send a check to support investigating specific research questions. The paywall deceives; science is neither an institut…
Most authors of scientific papers will gladly send you a free PDF of their papers if you ask them (assuming they remember to check their e-mail and respond in the first place). The profitability of their publisher is of no concern to them.
Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate
#158Earlier quoted context omitted.
>- Globalisation and the rule of law. Mostly that's a good thing. But SciHub would unlikely survive if Russia was unable to give US courts the middle finger regularly over the last decade. I am not convinced that the benefits of totalitarian regime outweighs the downsides but it is a thing I can't wait for 2030 when China overtakes the US as the worlds largest economy. We can all then be banned from the internet for…
That's not going to be a problem. When it comes to censorship of the Internet, your primary risk is the 5 / 14 / 19 eyes groups, not China. China isn't a critical part of the global Internet today and they'll be even less a part of it in another decade. They operate their own separate network that only poorly connects to the Internet, by design. That separation will increase considerably over this decade. Xi is curre…
I am happy to be corrected but interested at the counter point to such doom and gloom
Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate
#159Earlier quoted context omitted.
Why? The authors write the papers for free, the peer review is done by other scientists for free. Why should this "netflix for science" get to reap the profits by locking it behind a paywall? The reason why predatory publishers still exist is a coordination problem. The journals have prestige built up historically, and the scientists need to publish in prestigious journals for their career. It's a chicken and egg pro…
I'm not sure about how fair the whole system is, but if we assume the papers should be free, then a service where the papers are available is still reasonable to have a small fee. Someone needs to rent and service the servers, update the software, bandwidth costs money, etc..
Scientific papers use approximately zero bandwidth.
As a point of reference, if I hosted such a website on my home internet connection I would be paying approximately 0.00013 cents per upload. Even users downloading millions of papers would cost me less than a coffee.
Setting up and managing a proportionate payment system would cost more than just eating the bandwidth costs would.
Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate
#160Earlier quoted context omitted.
Seems like a large part of what's needed is just being able to make the pdfs machine-readable, by making decent plain-text versions of the text content. Right now, IIRC, there's no hands-off way to get the text of a pdf. Especially if there's weirdness like multiple-columns (sometimes happens with this stuff).
Currently working in this field and this is actually the cutting-edge(!!) but it will be 100% possible/robust within the next year or so I believe. Really cool ML techniques being used for htis.