Live data from Hacker News

Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

twitter.com

131–140 of 171 posts

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#131

Torrent seeding effort: https://www.reddit.com/r/DataHoarder/comments/nc27fv/rescue_... All papers on sci-hub are available as torrents from library genesis. The full collection contains 85 million articles (before this announcement), and is about 80TB. If anything ever happens to sci-hub or library genesis, there's enough people out there with backups that a replacement can be set up fairly quickly, albeit without t…

What we really need is an index of these torrents by DOI, and then ultimately by journal and issue. Are you aware of any work to make this happen?

The only one I'm aware of currently is https://github.com/sci-hub-p2p/sci-hub-p2p. Library genesis also hosts database dumps at https://libgen.rs/dbdumps/.

There's really a need though for more developers to get involved with building tools for more easily searching and working with the collection, ideally with a nice UI and integration with things like crossref. This is a massively valuable data set and it would be great to see what people can come up with. Lots of awesome potential for data mining too.

If it weren't for the legal issues (publishers using copyright law to restrict access to literature they got for free since they never pay authors for their work), there's no shortage of projects that could utilize this data and be enormously beneficial for the scientific community and humanity in general. Unfortunately such work can only be done in the shadows right now, which greatly limits the number of people/institutions likely to do so.

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#132

Earlier quoted context omitted.

What about a Netflix but for science information . Pay $15 a month for access to a rolling catalogue of science info.

You can more or less do this with scite's Citation Statement Search[1,2], or by setting email notifications when we detect new citation statements to one or more papers (grouped by a topic you're interested in like a disease or drug, an author, or more)[3]. Quick video of our citation statement search to give you a glimpse: https://www.youtube.com/watch?v=JYjCn-4uMJk [1] https://scite.ai [2] https://citation.to [3] h…

This self promotion seems off topic and unrelated to the person you are replying to.

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#133
post #5

I really hope sci-hub survives this. Sci-hub and libgen are like an entirely different internet, one allowing you to dive as deep as you wish into any technical subject. There’s really no comparison I have found anywhere for the depth of material available. People always point to Wikipedia, but all of that is surface level. If you to build something, research something, or just really delve into it, there’s no substi…

I really wish someone could build a better UI for this research internet. Hyperlinks for all references would be a good start. Finding some way to make some automatic glossary of definitions of technical terms would make scientific papers substantially more accessible too.

Seems like a large part of what's needed is just being able to make the pdfs machine-readable, by making decent plain-text versions of the text content. Right now, IIRC, there's no hands-off way to get the text of a pdf. Especially if there's weirdness like multiple-columns (sometimes happens with this stuff).

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#134

Earlier quoted context omitted.

Hyperlinking references in PDFs is trivial for authors if they use LaTeX and the journal/conference template supports it. It's just a matter of ensuring that the bibtex entry has an URL or a DOI, and most bibtex entries copied and pasted from curated sources already have them. If you are finding many papers without hyperlinked references, it's probably just because they're published in journals whose templates don't…

It's not adding the links that is hard. It's choosing the destination URL. Where do you link to? I guess you could link to doi.org. Probably better than nothing but still not ideal because it doesn't actually take you to the PDF. Can you show an example paper with links?

Linking directly to the PDF is usually the wrong choice. When you find a new paper, you often want to get the citation metadata, which the PDF document rarely contains in a convenient form. There are often multiple versions of the same paper, and you may want to determine which version you managed to find. Is it a preprint, the final authors' version, the published journal paper, an early version published in conference proceedings, or an unpublished extended version of the paper?

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#135
post #5

I really hope sci-hub survives this. Sci-hub and libgen are like an entirely different internet, one allowing you to dive as deep as you wish into any technical subject. There’s really no comparison I have found anywhere for the depth of material available. People always point to Wikipedia, but all of that is surface level. If you to build something, research something, or just really delve into it, there’s no substi…

That's the best part about a good idea once it's out there! It's hard to kill. Really wish we had come up with an alternative to 20 streaming sites...

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#136
post #5

I really hope sci-hub survives this. Sci-hub and libgen are like an entirely different internet, one allowing you to dive as deep as you wish into any technical subject. There’s really no comparison I have found anywhere for the depth of material available. People always point to Wikipedia, but all of that is surface level. If you to build something, research something, or just really delve into it, there’s no substi…

That's the best part about a good idea once it's out there! It's hard to kill. Really wish we had come up with an alternative to 20 streaming sites...

There is. Torrents. The UX can be amazing if you know how, but we dont want spoil the party by sharing.

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#137

Earlier quoted context omitted.

I really wish someone could build a better UI for this research internet. Hyperlinks for all references would be a good start. Finding some way to make some automatic glossary of definitions of technical terms would make scientific papers substantially more accessible too.

Seems like a large part of what's needed is just being able to make the pdfs machine-readable, by making decent plain-text versions of the text content. Right now, IIRC, there's no hands-off way to get the text of a pdf. Especially if there's weirdness like multiple-columns (sometimes happens with this stuff).

Currently working in this field and this is actually the cutting-edge(!!) but it will be 100% possible/robust within the next year or so I believe. Really cool ML techniques being used for htis.

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#138

Earlier quoted context omitted.

I really wish someone could build a better UI for this research internet. Hyperlinks for all references would be a good start. Finding some way to make some automatic glossary of definitions of technical terms would make scientific papers substantially more accessible too.

Seems like a large part of what's needed is just being able to make the pdfs machine-readable, by making decent plain-text versions of the text content. Right now, IIRC, there's no hands-off way to get the text of a pdf. Especially if there's weirdness like multiple-columns (sometimes happens with this stuff).

Like you said the hard parts are the unstructured data/images/tables. There are pretty-good(80% of the way there) solutions tho. But nothing that could handle millions of paper without error

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#139
post #29

Torrent seeding effort: https://www.reddit.com/r/DataHoarder/comments/nc27fv/rescue_... All papers on sci-hub are available as torrents from library genesis. The full collection contains 85 million articles (before this announcement), and is about 80TB. If anything ever happens to sci-hub or library genesis, there's enough people out there with backups that a replacement can be set up fairly quickly, albeit without t…

Is there any legal risk to users in North America that do this? Is this copyrighted material?

From watching the Aaron Schwartz documentary, hopefully once scihub makes these publishers obsolete they will no longer have the power to press charges. But yeah they can and will get the Feds involved and will ruin your life like Schwartz
Post reply on HN