Live data from Hacker News

Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

twitter.com

51–60 of 171 posts

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#51
True question: is there a value that elsevier&co are providing? Proofreading or Selecting the articles, by example? If so, even if their price reflect more their de facto monopoly than this value, if we want to replace them we need to also find a way to replace it. Could we outcompete them?

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#52
post #31

Earlier quoted context omitted.

Is it compressible or already compressed?

It's all PDF files, which have their own compression, so it's unlikely there would be substantial gain from additional compression. Each torrent has 100 zip files, and each zip file has 1000 PDFs, but the files are stored uncompressed within the zips (i.e. using the STORE method).

Is there some kind of searchable index included so that you can locate an article in a particular Zip? I'm assuming each article has some kind of ID numbers and the Zips are divided by ID range or something?

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#53
post #31

Earlier quoted context omitted.

Is it compressible or already compressed?

It's all PDF files, which have their own compression, so it's unlikely there would be substantial gain from additional compression. Each torrent has 100 zip files, and each zip file has 1000 PDFs, but the files are stored uncompressed within the zips (i.e. using the STORE method).

> It's all PDF files, which have their own compression, so it's unlikely there would be substantial gain from additional compression.

You could write a custom compressor that decompiles journal PDFs to valid TeX, then compresses that.

Or at the simpler end of what's technologically possible, you could at least extract shared assets such as fonts that appear in multiple files. Keep files from the same journal together to find more overlaps.

I suspect there's quite a large gain to be had from further compression, at least theoretically. Even more if you could accept some level of non-semantic loss.

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#54
post #45
post #10

In https://twitter.com/ringo_ring/status/1414342378765307907 she jokingly asks where her Nobel Peace Prize is. Given how often that prize is given out as a bully pulpit to advance a cause, and given the global debt that science owes to her, I think she really does deserve one.

She needs to step it up by killing thousands of innocent people from drone strikes to increase her chances.

Or start a war with her own citizens and use hunger as a weapon by destroying infrastructure, farming equipment and blocking foreign aid, creating a famine among millions.

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#55

Earlier quoted context omitted.

It's all PDF files, which have their own compression, so it's unlikely there would be substantial gain from additional compression. Each torrent has 100 zip files, and each zip file has 1000 PDFs, but the files are stored uncompressed within the zips (i.e. using the STORE method).

Is there some kind of searchable index included so that you can locate an article in a particular Zip? I'm assuming each article has some kind of ID numbers and the Zips are divided by ID range or something?

Yep! See https://github.com/sci-hub-p2p/artifacts/releases/tag/0

This project is in it's early stages and the documentation has quite some way to go, but the index that's part of the release contains all the necessary information. This tool also contains the code necessary to produce the index files if you have a local copy of the zips.

Each torrent contains 100,000 files, comprised of 100 zip files with 1,000 PDFs each. They are named by DOI. There's a database dump at (http://libgen.rs/dbdumps/) (scimag.sql.gz) which has the id -> DOI mapping and other information. The specific torrent and zip file can be determined based on the id; torrent = id/100000 and zip = id/1000.

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#56

Earlier quoted context omitted.

It's all PDF files, which have their own compression, so it's unlikely there would be substantial gain from additional compression. Each torrent has 100 zip files, and each zip file has 1000 PDFs, but the files are stored uncompressed within the zips (i.e. using the STORE method).

Is there some kind of searchable index included so that you can locate an article in a particular Zip? I'm assuming each article has some kind of ID numbers and the Zips are divided by ID range or something?

Sci-Hub database/index is available here: http://libgen.rs/dbdumps/scimag.sql.gz

and database documentation is available here: https://gitlab.com/lucidhack/knowl/-/wikis/References/Libgen...

also see introduction to Sci-Hub for developers: https://www.reddit.com/r/scihub/comments/nh5dbu/a_brief_intr...

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#57
post #31

Earlier quoted context omitted.

Is it compressible or already compressed?

It's all PDF files, which have their own compression, so it's unlikely there would be substantial gain from additional compression. Each torrent has 100 zip files, and each zip file has 1000 PDFs, but the files are stored uncompressed within the zips (i.e. using the STORE method).

But each PDF is compressed individually. The textual content of the papers must have a lot of redundancy between them, maybe there is some gain to get there?

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#58

Context: Sci-Hub stopped uploading new papers in December 2020, after being ordered to do so by an Indian court. There was (is?) some hope of winning the case which could make Sci-Hub legal in India. https://www.reddit.com/r/scihub/comments/mk46x4/scihub_v_els... https://news.ycombinator.com/item?id=26264378

[deleted]

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#59
post #5

I really hope sci-hub survives this. Sci-hub and libgen are like an entirely different internet, one allowing you to dive as deep as you wish into any technical subject. There’s really no comparison I have found anywhere for the depth of material available. People always point to Wikipedia, but all of that is surface level. If you to build something, research something, or just really delve into it, there’s no substi…

What about a Netflix but for science information . Pay $15 a month for access to a rolling catalogue of science info.

Why? The authors write the papers for free, the peer review is done by other scientists for free. Why should this "netflix for science" get to reap the profits by locking it behind a paywall? The reason why predatory publishers still exist is a coordination problem. The journals have prestige built up historically, and the scientists need to publish in prestigious journals for their career. It's a chicken and egg problem.

Re: Today Sci-Hub is 10 years old. I'll publish 2M new articles to celebrate

#60
post #44
post #29

Earlier quoted context omitted.

Is there any legal risk to users in North America that do this? Is this copyrighted material?

I encourage you to do this in the most audacious, and visible way possible. Let see how will they sue millions, upon millions of people. Make them face a fait accompli. They already lost.

[deleted]
Post reply on HN