Live data from Hacker News

Library-managed 'arXiv' spreads scientific advances rapidly and worldwide

ezramagazine.cornell.edu

61–70 of 141 posts

Re: Library-managed 'arXiv' spreads scientific advances rapidly and worldwide

#61

Ah, the original http://xxx.lanl.gov/ that I knew and loved in the 90's, when people thought we were surfing nudies in the Physics department and not papers on differential geometry. I helped establish and run the za.arxiv.org mirror at WITS University, mostly to learn how to configure RedHat, Apache, rsync and other tools. I'm glad it still exists.

I've worked at Los Alamos for 6 years and I didn't know this existed. Pretty cool.

Re: Library-managed 'arXiv' spreads scientific advances rapidly and worldwide

#62

Can I just make a general plea? You should upload your paper to arXiv. When you do, please upload your source (tex, or word I imagine), as well as a PDF. For the blind, PDF is the worst possible format, and tex and word are the best formats. Don't hide, or lose, the blind-accessible version of your paper.

When I worked at the Cornell Theory Center in the late 90s, I didn't know about the arXiv, but I sure as hell knew about screen readers. One of the best FORTRAN programmers (numerical analysis/applied math research faculty I think) was blind as a bat.

I learned a lot about not underestimating people there.

It's a strange university but it always makes me happy to think that it (specifically Ginsparg, and to some degree strogatz and the law scholars) pushed forward what the Web was supposed to be.

Not a place to buy stuff, but a place to learn stuff, without waiting for it to make its way through a hyper politicized review process so that it could be printed in a 17th century fashion and mailed to some corners of the world, eventually perhaps reaching a fraction of the people who could use it. Rather, everyone everywhere with a connection.

Anyone who doesn't deposit preprints (arXiv, biorxiv, or wherever) or who doesn't agitate for their coauthors to do so is not really in it for the science. It's fine to be competitive -- deposit yours first. Make the fucking discovery instead of parasitically piggybacking on those who do the work.

But that last part, that's hard. Very hard. As long as there is enough money left in academia to encourage lazy shits, those of us who care about scholarship will have to push, hard, to remove the last refuge of these scoundrels.

You're either on the side of justice -- open data, open formats, open scholarship -- or you are tacitly endorsing Elsevier & Springer, who haven't the slightest problem using crap like incremental JavaScript & mangled PDFs to deny access to scholarship even to those who have paid.

David (blocking on his last name) at CTC made that choice for me. He set an example that forced me to admit what was right. I hope others will do the same. It's the right thing to do.

Re: Library-managed 'arXiv' spreads scientific advances rapidly and worldwide

#63
post #47
post #46

>Eleven years ago Ginsparg joined the Cornell faculty, bringing what is now known as arXiv.org with him. (Pronounce it "archive." The X represents the Greek letter chi.) Been pronouncing it "ar ziv" until now. :P

But, "archive dot org" is an entirely different and also noteworthy organization!

I think that's "ark ive" where this is "ar chive". I'm gonna call it ar chive dot org now and get more stupid looks, aren't I? :/

Re: Library-managed 'arXiv' spreads scientific advances rapidly and worldwide

#64
post #46

>Eleven years ago Ginsparg joined the Cornell faculty, bringing what is now known as arXiv.org with him. (Pronounce it "archive." The X represents the Greek letter chi.) Been pronouncing it "ar ziv" until now. :P

It may be interesting to note that the ancient greek 'X' is suposed to be aspirated so the English pronunciation of 'chi' is almost certainly incorrect anyways.

Re: Library-managed 'arXiv' spreads scientific advances rapidly and worldwide

#65
post #19
post #14

Earlier quoted context omitted.

The flow is significantly slower than is required for checking work. I'm fairly sure my wifes papers didn't really take nearly two years to review. Machine learning is a field that's moving quickly and can also largely be checked quickly by other people in the field too. The problems that can't be checked easily are also those that won't be checked by the journals anyway (no journal I know of will retrain one of goog…

If you think peer-reviewing is slow then it's probably because you don't get the PEER part of the reviewing process. The PEERS are simply colleagues who have their own lives, priorities, and research. They don't stop what they are doing to review your wife's papers. And please, I see what's happening here: I have been portrayed as the "traditional journal apologist" although it's definitely what I wanted here. I simp…

> If you think peer-reviewing is slow then it's probably because you don't get the PEER part of the reviewing process.

I understand the process fully. It doesn't make it fast, nor does it make it a necessary cost to pay. Delaying access to content for several years does not solve a problem. A not-insignificant time was spent bouncing between people to sort out who was paying for the costs, then there is also a delay between acceptance and publication. This now averages just a month in pubmed, but papers can bounce around this point for a lot longer.

I am not arguing for pre-prints to replace traditional publishing, but the speed of spreading information is undeniably faster, and that's what started this whole chain of comments.

> Anyway, is there a study that shows how many pre-printed articles in ML have been retracted or refuted so far? Until this is done, and shown as small as in peer-reviewed journals, keep a small basket.

I don't get the phrase "keep a small basket", but no I've not seen this. The point is simply that PEERS (if we need to shout the word) can in many cases replicate the work and assess the results much more quickly than the traditional review & edit process. I picked the field because I see people re-implement work described in the papers very quickly, or the original authors share the trained models and code.

I'd also caution against using "retracted" papers as a measure, some journals charge for a retraction.

Re: Library-managed 'arXiv' spreads scientific advances rapidly and worldwide

#66

This article misses one of the biggest value-adds of arXiv, at least in my field (Statistics): since almost everyone posts to arXiv, you can almost always find a free version of a published and potentially pay-walled paper. In the past, publishing in a peer-reviewed journal would (1) improve the paper through peer review, (2) signal the quality of the paper based on the prestige of the journal, and (3) distribute the…

Publishing sometimes does 1), and rarely does 2) (as a statistician you surely know that the relationship between impact factor and retraction is nonlinear and rises in strength as you get into CNS, NEJM, and the like).

I review for others because others have done 1) for me. But I'll never review for Elsevier, and lately I've had the luxury of reviewing for the most cited of open journals (by operating bioRxiv, and accepting direct submissions from it, I claim that Genome Research is "close enough").

It makes me very happy that this is possible (my CV has not suffered for only publishing as first author, and whenever possible as senior or co-senior, in fully open journals). I'm pretty sure this wasn't possible for most people a few short years ago. That engenders optimism about the future of scholarship, for me at least.

Hopefully you as well.

Re: Library-managed 'arXiv' spreads scientific advances rapidly and worldwide

#67
post #38
post #35

Earlier quoted context omitted.

As a scientist, give me the data and I will know what to do with it. AFAIK in the http://biorxiv.org/ they provide some statistics and it does not represent a problem.

As a fellow scientist, I'm much more concerned with how others will interpret these access data. I'm not excited about the prospect of yet another unreliable signal for e.g. hiring committees to latch onto, as they often do with journal impact factors and such. It might be nice if ArXiv would perhaps provide the data to researchers on request. Just curious -- what kinds of questions would you use this data to answer?

I want the data for the same reasons that any content producer in the Internet wants it. Bloggers, youtubers, any company...everyone. Despite the noise this data might contain, it seems it's useful for everyone except for scientists...to whom I am surprised to hear that it's better not to give the data, in case they misinterpret it. Very risky statement and precedent.

Re: Library-managed 'arXiv' spreads scientific advances rapidly and worldwide

#68
post #32

I hate arXiv, I can never figure out where is the PDF, if there is a PDF... long live eprint.

I'm not sure whether you're serious, but on any article's page there a "Download" section with a link to the PDF (labelled "PDF").

Not always.

Re: Library-managed 'arXiv' spreads scientific advances rapidly and worldwide

#69
post #59
post #51

Earlier quoted context omitted.

Besides disputes involving patents, papers that are withing a short time frame of each other are usually understood to be cases of parallel invention, it's happening quite frequently in deep learning atm since there are still a lot of relatively low hanging ideas, to the extent that people are commenting/joking that they consider the risk of colliding with someone else when deciding what to work on.

> withing a short time frame of each other are usually understood to be cases of parallel invention I see what you are saying, but I don't think it's that cut and dry, otherwise I could just take someone else's work from yesterday (or whatever a short time frame is), and re-solve it (easily -since now the tricky parts have been revealed) and post it today - tada, I parallel invented it!

Typically the work in a paper, if substantial, is done over a long time, so even if the main destination ends up being same, it's unlikely the route and sidestops are the same. So often you can wriggle a little bit and expand the paper sideways, so that it is still publishable work even if the other work is given priority.

It's actually not that rare to have similar papers appear in arXiv one or two weeks later after you submit --- to me, it happened several times within last few years. In these cases, it is possible to see that the approach differs enough (and moreover, often you know the people in question, or you know someone who does).

Re: Library-managed 'arXiv' spreads scientific advances rapidly and worldwide

#70

Can I just make a general plea? You should upload your paper to arXiv. When you do, please upload your source (tex, or word I imagine), as well as a PDF. For the blind, PDF is the worst possible format, and tex and word are the best formats. Don't hide, or lose, the blind-accessible version of your paper.

PDF readers manage to parse text back somehow effectively, maybe not on formula / formatting heavy PDFs. Anyway good call, accessibility is not only for mainstream websites. I'm sure the blind dude that aced math classes in college would agree.
Post reply on HN