Live data from Hacker News

Just 11% of 53 cancer research papers were reproducible

nature.com

51–60 of 184 posts

Re: Just 11% of 53 cancer research papers were reproducible

#51

Perhaps we are looking for the wrong measure of success. This is not inclined plane ball rolling in fifth period Physics with Mr. Johannes. Reproducibility in cutting edge experiment is success. It means that what is being measured - even if the wrong thing - is within the control of the experimenter. It is absurd, but practical, to publish experimental results with the implication that they are reproducible based on…

I wonder if it would be possible to produce a journal that only released an issue when it has, say, 10 articles? No monthly/quarterly schedule, you could have an issue release next week, or next year.

Re: Just 11% of 53 cancer research papers were reproducible

#52
post #7

After hearing similar things about psychology papers, this is rather disconcerting. This is why I am a climate change skeptic, I don't know whether it is happening due to CO2 or not, but I am confident that science can't be confident when they can't do control experiments. You can't control for any variable when it comes to the climate, let alone all the reasonable ones. We have infinitely more capabilities to contro…

The possibility of CO2 causing global warming isn't 100% but it isn't very low either. It is the risk we are talking about, the risk of doing nothing and let the man made green house (if there is one) turn earth into an irreversible disastrous environment. Most CO2 emitting energy sources are not sustainable anyway and many of them emit other proved pollution as well. There is really nothing lost to going green energ…

First, thanks for trying to add to the discussion, I appreciate it. I disagree on two points:

1. I make no judgement about whether or not CO2 causes global warming. It may cause it, it may cause it but to a lesser (or greater) effect than sun cycles, cosmic radiation (and it's effect on cloud formation) or natural variation in the climate. I'm just asserting that my confidence in knowing one way or the other when you can't create control experiments is low. Ultimately proving causation instead of simple correlation is something that I believe is near impossible for something as large and complicated as the climate. Especially when the data being analysed for correlation relies hundreds of thousands of years of extrapolated proxy data (tree rings, ice cores, sediment, etc).

2. (this is tangential to the discussion of reproducibility in science) I think there is something lost to going to green energy - cost and power density. The unfortunate fact is that fossil fuels are incredibly cheap compared to green energy and they are more easily transportable. Carbon regimes threaten the ability of the third world to build itself out of poverty and threaten the global economy in general. I'm not saying we shouldn't invest in new forms of energy (especially wind and nuclear), in fact I think we certainly should. But I don't want to hamstring economic progress for some calamitous event with something I have low confidence will actually happen.

Re: Just 11% of 53 cancer research papers were reproducible

#53
There are many reasons so many papers are not reproducible.

1.) The prospective data collection methods used in a lot of studies are deeply flawed and worse aren't documented.

2.) The retrospective data used in many studies is poorly validated and the quality is rarely a concern for most universities.

3.) The people who publish these papers are very intelligent, most far more intelligent than I am, but you'd be surprised how many of them don't have a very good handle on basic statistics. Not because they aren't smart enough to learn but mostly because they have no interest in it.

4.) There is a lot of pressure put on researchers to publish papers even if they aren't ready and as a result mistakes are made, there really needs to be less of an emphasis on the quantity of papers that are published and more emphasis on quality.

For anyone who might be interested in such things, if you ever find yourself locked in a dungeon with nothing but a computer loaded with every journal article ever published you can find articles involving similar cohorts from the exact same universities that have different results.

Re: Just 11% of 53 cancer research papers were reproducible

#54
post #49

What is more troubling is not that so few of these results results are reproducible, but that it appears almost no-one is trying to reproduce the results of earlier studies. The ability of the scientists who wrote the paper to even get access to the resources necessary to try and reproduce the results is limited. I'm reminded of Feymann's Cargo Cult Science: "I was shocked to hear of an experiment done at the big acc…

In practical terms, reproducing someone else's work most of the time boils down to redoing someone else's PhD thesis. Which is both not very interesting and doesn't help getting your own research done.

Re: Just 11% of 53 cancer research papers were reproducible

#55
post #23

Wow, that's sad. Are we really so blind when it comes to cancer research? I recall that oncology journals usually have a ludicrously high impact index (as ludicrously as 5 digit IIRC); that means there are a lot of citations, which is an indicator that there is a lot of research going on. And with a lot of research going on, well, you can expect a lot of false positives. So, I'm wondering, could this be a case of che…

"And with a lot of research going on, well, you can expect a lot of false positives." Why would you expect a lot of false positives? Aren't results supposed to be reproducible at least 95% of the time?

I was thinking about absolute numbers, so 5% of a lot can still be a lot.

Re: Just 11% of 53 cancer research papers were reproducible

#56
post #4

This is one of the reasons I went into math and software. In math there is only black and white, provably true and patently wrong. Still, part of the reason this is so is that to show any results everything must be laid on the table. CS and Math papers are therefore open by design and the world has benefited. I hope the rest of science will move in this direction.

> In math there is only black and white, provably true and patently wrong.

You'd probably be interested in Godel's incompleteness theorem.

Also, exact reproduction of other people's experimental work is highly unusual in CS, despite the theoretical possibility. Usually the materials and methods sections of papers in hard sciences like physics, chemistry, and biology are much more detailed, to the point of being recipes.

Finally, CS is not really a science. It lies somewhere between math and engineering, which are also not sciences.

Re: Just 11% of 53 cancer research papers were reproducible

#57
This editorial commentary and the article on which it is based are part of an ongoing effort to improve the quality of scientific publication in a number of disciplines. The Retraction Watch group blog

http://retractionwatch.wordpress.com/

by two experienced science journalists picks up many--but not all--of the cases of peer-reviewed research papers being retracted later from science journals.

Psychology as a discipline has been especially stung by papers that cannot be reproduced and indeed in many cases have simply been made up.

http://www.nytimes.com/2013/04/28/magazine/diederik-stapels-...

That has prompted statistically astute psychologists such as Jelte Wicherts

http://wicherts.socsci.uva.nl/

and Uri Simonsohn

http://opim.wharton.upenn.edu/~uws/

to call for better general research standards that can be practiced as checklists by researchers and journal editors so that errors are prevented.

Jelte Wicherts writing in Frontiers of Computational Neuroscience (an open-access journal) provides a set of general suggestions

Jelte M. Wicherts, Rogier A. Kievit, Marjan Bakker and Denny Borsboom. Letting the daylight in: reviewing the reviewers and other ways to maximize transparency in science. Front. Comput. Neurosci., 03 April 2012 doi: 10.3389/fncom.2012.00020

http://www.frontiersin.org/Computational_Neuroscience/10.338...

on how to make the peer-review process in scientific publishing more reliable. Wicherts does a lot of research on this issue to try to reduce the number of dubious publications in his main discipline, the psychology of human intelligence.

"With the emergence of online publishing, opportunities to maximize transparency of scientific research have grown considerably. However, these possibilities are still only marginally used. We argue for the implementation of (1) peer-reviewed peer review, (2) transparent editorial hierarchies, and (3) online data publication. First, peer-reviewed peer review entails a community-wide review system in which reviews are published online and rated by peers. This ensures accountability of reviewers, thereby increasing academic quality of reviews. Second, reviewers who write many highly regarded reviews may move to higher editorial positions. Third, online publication of data ensures the possibility of independent verification of inferential claims in published papers. This counters statistical errors and overly positive reporting of statistical results. We illustrate the benefits of these strategies by discussing an example in which the classical publication system has gone awry, namely controversial IQ research. We argue that this case would have likely been avoided using more transparent publication practices. We argue that the proposed system leads to better reviews, meritocratic editorial hierarchies, and a higher degree of replicability of statistical analyses."

Uri Simonsohn provides an abstract (which links to a full, free download of a funny, thought-provoking paper)

http://papers.ssrn.com/sol3/papers.cfm?abstract_id=2160588

with a "twenty-one word solution" to some of the practices most likely to make psychology research papers unreliable. He has a whole site devoted to avoiding "p-hacking,"

http://www.p-curve.com/

an all too common practice in science that can be detected by statistical tests. He also has a paper posted just a few days ago

http://papers.ssrn.com/sol3/papers.cfm?abstract_id=2259879

on evaluating replication results (the issue discussed in the commentary submitted to open this thread) with more specific tips on that issue.

"Abstract: "When does a replication attempt fail? The most common standard is: when it obtains p>.05. I begin here by evaluating this standard in the context of three published replication attempts, involving investigations of the embodiment of morality, the endowment effect, and weather effects on life satisfaction, concluding the standard has unacceptable problems. I then describe similarly unacceptable problems associated with standards that rely on effect-size comparisons between original and replication results. Finally, I propose a new standard: Replication attempts fail when their results indicate that the effect, if it exists at all, is too small to have been detected by the original study. This new standard (1) circumvents the problems associated with existing standards, (2) arrives at intuitively compelling interpretations of existing replication results, and (3) suggests a simple sample size requirement for replication attempts: 2.5 times the original sample."

The writers of scientific papers have a responsibility to do better. And the readers of scientific papers that haven't been replicated (or, worse, press releases about findings that haven't even been published yet) also have a responsibility not to be too credulous. That's why my all-time favorite link to share in comments on HN is the essay "Warning Signs in Experimental Design and Interpretation" by Peter Norvig, LISP hacker and director of research at Google, on how to interpret scientific research.

http://norvig.com/experiment-design.html

Check each submission to Hacker News you read for how many of the important issues in interpreting research are NOT discussed in the submission.

Re: Just 11% of 53 cancer research papers were reproducible

#58
post #48
post #23

Wow, that's sad. Are we really so blind when it comes to cancer research? I recall that oncology journals usually have a ludicrously high impact index (as ludicrously as 5 digit IIRC); that means there are a lot of citations, which is an indicator that there is a lot of research going on. And with a lot of research going on, well, you can expect a lot of false positives. So, I'm wondering, could this be a case of che…

But who is going to pay to reproduce that research? What if that research took years to do?

Why do the research in the first place if the only outcome is a number, floating in space, disconnected from anything and unreproducible?

You're still sort of operating on the idea that the unreproducible papers have some sort of abstract value to them, and therefore we shouldn't slow the flow of them lest we ruin their value. But they don't. They're worthless. They're worse than worthless. They're of negative value. It would be far better to slow down and verify that what we think we know is actually true, because in the end that would actually both faster and a more efficient use of resources. Basically, instead of learning worse than nothing (thinking we know something but actually being wrong), we'd learn something. That's a pretty decent upgrade.

Re: Just 11% of 53 cancer research papers were reproducible

#59
I work in biomedical research, and this finding has been discussed quite broadly. Most researchers don't believe it.

Exactly reproducing novel findings usually requires a significant investment into the underlying procedures, which most places will not undertake. Reproduction instead usually occurs as part of an extension of the initial findings. The complexity of biology means that the first findings are frequently not reproduced exactly the same way as the original, but this does not detract from the "direction" of the initial findings.

For example, the first finding might be that protein A promotes tumor growth by modifying protein B. An extension of these findings might be Protein A sometimes modifies protein B, and when it does tumor growth is stimulated, but it mostly does not modify protein B, it instead modifies protein C, which suppresses tumor growth. In this case, were the first results replicated? Yes, and no.

This is how most biomedical research proceeds...

Re: Just 11% of 53 cancer research papers were reproducible

#60
post #54
post #49

What is more troubling is not that so few of these results results are reproducible, but that it appears almost no-one is trying to reproduce the results of earlier studies. The ability of the scientists who wrote the paper to even get access to the resources necessary to try and reproduce the results is limited. I'm reminded of Feymann's Cargo Cult Science: "I was shocked to hear of an experiment done at the big acc…

In practical terms, reproducing someone else's work most of the time boils down to redoing someone else's PhD thesis. Which is both not very interesting and doesn't help getting your own research done.

I agree that reproducing someone else's experiment won't help you get your PhD closer to completion (because you should be doing your own experiments and publishing no matter what, if you expect to land that lecturing position, i.e. publish or perish).

Reproducibility is one of the main principles of the scientific method (those papers shouldn't even been accepted if they don't contain enough information about how to reproduce the experiments they describe)

If an experiment is so complex that it can be compared to redoing someone else's thesis it can be either that the thesis is very simple or the experiment is so complex that it probably proves nothing.

Post reply on HN