Live data from Hacker News

Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

twitter.com

141–150 of 171 posts

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#141

Earlier quoted context omitted.

What we really need is an index of these torrents by DOI, and then ultimately by journal and issue. Are you aware of any work to make this happen?

The only one I'm aware of currently is https://github.com/sci-hub-p2p/sci-hub-p2p . Library genesis also hosts database dumps at https://libgen.rs/dbdumps/ . There's really a need though for more developers to get involved with building tools for more easily searching and working with the collection, ideally with a nice UI and integration with things like crossref. This is a massively valuable data set and it would b…

Isn't libgen already distributed via IPFS too?

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#142

Torrent seeding effort: https://www.reddit.com/r/DataHoarder/comments/nc27fv/rescue_... All papers on sci-hub are available as torrents from library genesis. The full collection contains 85 million articles (before this announcement), and is about 80TB. If anything ever happens to sci-hub or library genesis, there's enough people out there with backups that a replacement can be set up fairly quickly, albeit without t…

What we really need is an index of these torrents by DOI, and then ultimately by journal and issue. Are you aware of any work to make this happen?

I believe there is a database dump available with this information:

http://libgen.rs/dbdumps/scimag.sql.gz

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#143
post #5

I really hope sci-hub survives this. Sci-hub and libgen are like an entirely different internet, one allowing you to dive as deep as you wish into any technical subject. There’s really no comparison I have found anywhere for the depth of material available. People always point to Wikipedia, but all of that is surface level. If you to build something, research something, or just really delve into it, there’s no substi…

That's the best part about a good idea once it's out there! It's hard to kill. Really wish we had come up with an alternative to 20 streaming sites...

PeerTube at https://joinpeertube.org/ or Owncast at https://owncast.online/

Both support live streams, and are (being) federated and ever more integrated with other Fediverse apps.

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#144
post #10

In https://twitter.com/ringo_ring/status/1414342378765307907 she jokingly asks where her Nobel Peace Prize is. Given how often that prize is given out as a bully pulpit to advance a cause, and given the global debt that science owes to her, I think she really does deserve one.

For a woman in the similar age, I would say she is more deserving of Nobel Peace Prize than Malala Yousafzai. The risk that Malala takes in advocating for women right in Islamic countries is admirable, there is no denying in that. However, her impact are minuscule compared to Alexandra's in the big picture of progression as a human race. Malala's activism has not changed much on the course of women right in the count…

You do not need to uplift one impactful woman by putting down another.

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#145

Earlier quoted context omitted.

The only one I'm aware of currently is https://github.com/sci-hub-p2p/sci-hub-p2p . Library genesis also hosts database dumps at https://libgen.rs/dbdumps/ . There's really a need though for more developers to get involved with building tools for more easily searching and working with the collection, ideally with a nice UI and integration with things like crossref. This is a massively valuable data set and it would b…

Isn't libgen already distributed via IPFS too?

Yes, LibGen is mirrored on IPFS:

https://news.ycombinator.com/item?id=25209246

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#146

Earlier quoted context omitted.

Have you come across scite ( https://scite.ai ) yet? We're also innovating in this space by extracting citation statements from full-text articles and classifying their intent. So let's say paper A cites paper B. If you look at paper B, we show you: - how many times it was cited - the direct paragraphs from paper A where it was cited - the sections from paper A where paper B was referenced - ... and a lot more You ca…

No one cares. This thread is about free access to papers and not another paid service that forces you to pay monthly fees for something that could be a free service. In that sense you aren't any better than large online publishers. 8 bucks a month for a scientific paper search engine? Really?

[deleted]

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#147
post #122

I find it fascinating that SciHub seems to be highlighting two large issues. - Globalisation and the rule of law. Mostly that's a good thing. But SciHub would unlikely survive if Russia was unable to give US courts the middle finger regularly over the last decade. I am not convinced that the benefits of totalitarian regime outweighs the downsides but it is a thing - copyright law is not patent law, science is not pat…

>- Globalisation and the rule of law. Mostly that's a good thing. But SciHub would unlikely survive if Russia was unable to give US courts the middle finger regularly over the last decade. I am not convinced that the benefits of totalitarian regime outweighs the downsides but it is a thing I can't wait for 2030 when China overtakes the US as the worlds largest economy. We can all then be banned from the internet for…

That's not going to be a problem. When it comes to censorship of the Internet, your primary risk is the 5 / 14 / 19 eyes groups, not China.

China isn't a critical part of the global Internet today and they'll be even less a part of it in another decade. They operate their own separate network that only poorly connects to the Internet, by design. That separation will increase considerably over this decade.

Xi is currently putting new restraints into place to pull Chinese tech companies back even further from the Internet and into their own isolated network.

When China becomes the largest economy by GDP, it'll be meaningless to the operation of the Internet, which they'll only kinda-sorta be a part of.

Further, China is now widely regarded as the top adversary to the US and the West. That context will get increasingly confrontational and war-like in the coming years. Nearly all members of Congress are on board the anti-China bandwagon now, they've all gotten the message from above (the military industrial complex, which dictates nearly all foreign policy). The cultural atmosphere will increasingly become like it was when the USSR was the primary adversary for decades. As that confrontation increases, China's influence over the Internet will be intentionally reduced by the powers that actually do control the Internet today. China sees that coming as well and is taking steps ahead of time to reduce its exposure, points of influence and risk. At this point China views a military confrontation with the West as close to inevitable (which recent Xi speeches have elaborated on).

This increasing separation effort by China is in part designed to make it possible for China to attempt to destroy/damage the Internet - if it comes to that - without posing much terminal risk to their network and economy in the process. If they take down the Internet, it'll butcher the economies of their adversaries, while their own network remains highly functional. This is something the West is almost entirely unprepared for, and China is aggressively preparing for it; an epic mistake by the West.

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#148
I love Sci-Hub, even though I never use it. I just really like what it's doing for science. However, the law has failed us. It's very clear that Sci-Hub is overwhelmingly a Good Thing, yet governments want to tear it down because it's threatening companies' revenue streams too much.

Given how important it is, and how at risk it is, I think it's very important to find a technological solution to keep it up. We have the technology to distribute the papers (torrents) as well as a search index. I really hope that either Alexandra starts using these technologies more, or the technologies mature enough to be usable.

Then Sci-Hub would be unkillable.

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#149

Earlier quoted context omitted.

Have you come across scite ( https://scite.ai ) yet? We're also innovating in this space by extracting citation statements from full-text articles and classifying their intent. So let's say paper A cites paper B. If you look at paper B, we show you: - how many times it was cited - the direct paragraphs from paper A where it was cited - the sections from paper A where paper B was referenced - ... and a lot more You ca…

No one cares. This thread is about free access to papers and not another paid service that forces you to pay monthly fees for something that could be a free service. In that sense you aren't any better than large online publishers. 8 bucks a month for a scientific paper search engine? Really?

Hiya,

Well, I definitely agree with your sentiment in a normative sense that scientific papers should be free and readily accessible to all -- in part because a lot of it is funded through tax dollars!

But given the current state of affairs, we're looking at making that information accessible to people without having to pay exorbitant fees to access individual research. We also offer steep discounts for students or anyone in academia.

With that in mind I would push back a little that we're just a scientific paper search engine -- our system does a lot of work in extracting and classifying those citation statements, which makes it more powerful than traditional scientific search engines.

And besides just using our search, a huge time-saving value of our service is the report pages which helps you quickly build a qualitative understanding of how something was cited.

Even if all scientific papers were freely accessible, our report pages allow you to see the direct, relevant snippets from citing papers without having to manually read each and every single one. I think that is quite valuable!

I know I've gone on a little tangent from the original discussion about scihub, and having free and open access to papers, but I did just want to throw that in because I think it's an important distinction. And as much as we all want that free and open world to exist, I think it's also interesting to think about how we can open up that information for people in the interim.

Best,

Ashish

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#150

Earlier quoted context omitted.

Seems like a large part of what's needed is just being able to make the pdfs machine-readable, by making decent plain-text versions of the text content. Right now, IIRC, there's no hands-off way to get the text of a pdf. Especially if there's weirdness like multiple-columns (sometimes happens with this stuff).

Currently working in this field and this is actually the cutting-edge(!!) but it will be 100% possible/robust within the next year or so I believe. Really cool ML techniques being used for htis.

How well does it work with old OCR'd PDFs? :)
Post reply on HN