Live data from Hacker News

Why the world of scientific research needs to be disrupted

gigaom.com

51–60 of 86 posts

Re: Why the world of scientific research needs to be disrupted

#51
Actually, there are quite a few points where modern technologies would allow for a significant streamlining of research processes. Off the top of my head (as I work in the field):

1. data acquisition - you won't believe in how many labs devices with a computer interface (gpib or whatever) are running in standalone mode, costing hours in grad student hours where parameters are changed and values are read by hand - in the best case people duct-tape something together with a labview program. No. Just No.

2. collaborative data sharing - if you want to show your boss a graph, you email him a jpg - where is the site to upload a csv and show a graph to other people/edit together?

3. Writing papers: The state of the art is mailing a LaTeX(!) or Doc(!) file to your colleagues with .v1.edited appended... Reviewing the published material is just the last step.

PS: I'm working on a solution for #3 (Etherpad + LaTeX preview + export in the appropriate journal format). Drop me a line if you're interested in details.

Re: Why the world of scientific research needs to be disrupted

#53

Two comments: 1. Scientists publish in slow peer reviewed journals, because scientific publication has to be peer reviewed. The reviewers are a very small selection of other scientists, who are expert in the field of the publication. As long as the reviewers don't signal white smoke, a publication is no science by definition. This way of work is important to keep quality excellent. We certainly don't need more quickl…

The article makes at lest one important point. There are only two incentives for a scientist: publications and grant money. Getting published fast reduces quality of research. Also, the focus on publications hinders other research outcomes like publicly available data or software.

Industry or start-ups may not make full use of publications that attack the same problem over and over again but they surely can improve with a new data or newly implemented algorithms.

[EDIT] Publicly available data and software make replicability more realistic. Currently lack of details in publications make it almost impossible.

If only there was some way to influence NIH and NSF grant requirements ...

Re: Why the world of scientific research needs to be disrupted

#54
post #43
post #25

Earlier quoted context omitted.

>Speed is important for a number of factors, mistakes cost less if they can be quickly corrected. So being risky and publishing about your failures becomes useful (no one publishes failures...). Speed also reduces duplicate work. You're missing the point of scientific publication. The entire point of publication is so others can reproduce your work. Speed doesn't work if you need a billion data points taken over 20 y…

> You're missing the point of scientific publication. The entire point of publication is so others can reproduce your work. Speed doesn't work if you need a billion data points taken over 20 years to prove a long term issue. In theory, that's the point. In practice, no journals ever publish replications, so nobody wastes their time reproducing others' work when they could be working on something publishable or their…

>In practice, no journals ever publish replications, so nobody wastes their time reproducing others' work when they could be working on something publishable or their next grant proposal.

Uh, this is exactly how Science works. It's not worth publishing the exact same results of the exact same experiment by multiple people. That only adds noise to the discussion.

If I arrive to the same conclusion after running the experiment again, then there is little benefit to anyone if I do a full writeup and publish it. On the other hand, if I am unable to arrive to the same results, there is a tremendous value in publishing my findings. Was the original study flawed? Were my own methods? That's what peer review and publishing results help determine.

Re: Why the world of scientific research needs to be disrupted

#55

Two comments: 1. Scientists publish in slow peer reviewed journals, because scientific publication has to be peer reviewed. The reviewers are a very small selection of other scientists, who are expert in the field of the publication. As long as the reviewers don't signal white smoke, a publication is no science by definition. This way of work is important to keep quality excellent. We certainly don't need more quickl…

Why do we need the process of peer review?

Peer review is not robust against even low levels of collusion (http://arxiv.org/abs/1008.4324v1). Scientists who win the Nobel Prize find their other work suddenly being heavily cited (http://www.nature.com/news/2011/110506/full/news.2011.270.ht...), suggesting either that the community either badly failed in recognizing the work's true value or that they are now sucking up & attempting to look better by the halo effect. (A mathematician once told me that often, to boost a paper's acceptance chance, they would add citations to papers by the journal's editors - a practice that will surprise none familiar with Goodhart's law and the use of citations in tenure & grants.)

Physicist Michael Nielsen points out (http://michaelnielsen.org/blog/three-myths-about-scientific-...) that peer review is historically rare (just one of Einstein's 300 papers was peer reviewed! the famous _Nature_ did not institute peer review until 1967), has been poorly studied (http://jama.ama-assn.org/cgi/content/abstract/287/21/2784) & not shown to be effective, is nationally biased (http://jama.ama-assn.org/cgi/content/full/295/14/1675), erroneously rejects many historic discoveries (one study lists "34 Nobel Laureates whose awarded work was rejected by peer review" (http://www.canonicalscience.org/publications/canonicalscienc...); Horribin 1990 (http://jama.ama-assn.org/content/263/10/1438.abstract) lists others like the discovery of quarks), and catches only a small fraction (http://jama.ama-assn.org/cgi/content/abstract/280/3/237) of errors. And fraud, like the one we just saw in psychology? Forget about it (http://www.plosone.org/article/info%3Adoi%2F10.1371%2Fjourna...);

> "A pooled weighted average of 1.97% (N = 7, 95%CI: 0.86–4.45) of scientists admitted to have fabricated, falsified or modified data or results at least once –a serious form of misconduct by any standard– and up to 33.7% admitted other questionable research practices. In surveys asking about the behaviour of colleagues, admission rates were 14.12% (N = 12, 95% CI: 9.91–19.72) for falsification, and up to 72% for other questionable research practices....When these factors were controlled for, misconduct was reported more frequently by medical/pharmacological researchers than others."

No, peer review is not the secret sauce of science. Replication is more like it.

Re: Why the world of scientific research needs to be disrupted

#56
post #34

Earlier quoted context omitted.

> 1. Are capable of understanding everything in the paper Agreed. In my academic field (Debuggers), to understand and know enough to give a real review of the subject requires reading what I would estimate to be somewhere around a thousand papers. It also requires keeping up with the major industrial producers of the product. I think the field has something like two to three thousand papers right now, I haven't check…

Out of curiosity, is there some place to locate a clear listing of the seminal works in your field? I'm genuinely curious how much of the difficulty here is actual breadth/depth of the subject matter, and how much of the difficulty is due to some systemic inefficiency in the way research is published and consumed.

The trick is usually to look at the references section of a paper. If there's an old paper in there (say, >10 years old), then it's probably seminal. If you see the same paper in a lot of reference sections, then it's probably seminal.

So: to know what's seminal so you can skip reading a bunch of papers, you need to read a bunch of papers. Right.

Re: Why the world of scientific research needs to be disrupted

#57
post #40

Two comments: 1. Scientists publish in slow peer reviewed journals, because scientific publication has to be peer reviewed. The reviewers are a very small selection of other scientists, who are expert in the field of the publication. As long as the reviewers don't signal white smoke, a publication is no science by definition. This way of work is important to keep quality excellent. We certainly don't need more quickl…

I'm sorry, your claim seems to be: The process is working perfectly right now. This is absurd. The scientific method is a lot more open to interpretation than you suggest, and indeed it is interpreted with wide variety in different disciplines. The methods of science will be disrupted and improved. Bright lay people will absolutely be able to poke holes in research once the methods and data are completely open, as ha…

> Bright lay people will absolutely be able to poke holes in research

By the time they can astutely manage the actual field with understanding, they will no longer be lay people.

Re: Why the world of scientific research needs to be disrupted

#58
post #54
post #43

Earlier quoted context omitted.

> You're missing the point of scientific publication. The entire point of publication is so others can reproduce your work. Speed doesn't work if you need a billion data points taken over 20 years to prove a long term issue. In theory, that's the point. In practice, no journals ever publish replications, so nobody wastes their time reproducing others' work when they could be working on something publishable or their…

>In practice, no journals ever publish replications, so nobody wastes their time reproducing others' work when they could be working on something publishable or their next grant proposal. Uh, this is exactly how Science works. It's not worth publishing the exact same results of the exact same experiment by multiple people. That only adds noise to the discussion. If I arrive to the same conclusion after running the ex…

Not publishing positive replications is just as bad a problem as not publishing negative original results. We have things like BigTable and Hadoop now; if 100 laboratories repeat an experiment and publish their results, that just means we can raise our confidence in the result by the sum of their likelihood ratios.

Getting more data improves the accuracy your results even better than using more sophisticated algorithms: http://www.catonmat.net/blog/theorizing-from-data-by-peter-n...

Re: Why the world of scientific research needs to be disrupted

#59
post #20

Earlier quoted context omitted.

There's been an attempt to measure paper quality using the "impact factor", but as with seemingly every measure, it's led to shenanigans. Some publishers encourage authors to cite papers from their journals, and some editors write survey papers that happen to cite a bunch of articles from their journals. A recent AMS Notices covered this in math, and I sort of expect it's worse in other disciplines.

To me it looks a lot like a manipulation of search engine. Maybe Google or Microsoft will not disclose their rank inflation counter-tactics, but what they could do is to provide their solution for this specific verticals. As much as I'd like open-source effort for that, I understand that knowing the rules allows easy manipulation.

Well, there already is Google Scholar. I don't know precisely how much effort goes into the ranking algorithms used there -- but I do already use it in preference to Pubmed.

Re: Why the world of scientific research needs to be disrupted

#60
post #9

A very practical problem with making data available publicly is privacy. This is likely not going to be an issue for data coming out of the LHSC, but comes into play in CS research. I know a bunch of researchers who work with large datasets that have user locations, cell tower communication, social network data, etc. Of course, even researchers work with data that is annonymized. But almost nothing is truly anonymous…

"Researchers must also go through IRBs (independent review boards) at their institution prior to engaging in research that deals with human subjects. If the collected data is going to be made available publicly, it makes the process more arduous."

That is not really my experience. Getting IRB approval to release human data just means proper de-identification which one should do anyway. For example, generally subjects are given randomly generated IDs with only the PI having the master list.

Post reply on HN