Live data from Hacker News

The Irreproducibility Crisis of Modern Science

nas.org

251–260 of 265 posts

Re: The Irreproducibility Crisis of Modern Science

#251
post #229
post #228

Earlier quoted context omitted.

Then that is incorrect. The situation is simultaneously better and worse than your interpretation: > while a low P value indicates that your data are unlikely > assuming a true null, it can’t evaluate which of two > competing cases is more likely: > * The null is true but your sample was unusual. > * The null is false. You're conflating the two populations, and indeed we don't know which is which. But the P-value is…

> it can’t evaluate which of two competing cases is more likely Of course it can. In fact, it can tell you exactly how much more likely one case is versus the other. P=0.05 means that there is a 5% chance that the null is true and hence a 95% chance the the null is false. > says something about the likelihood of the _hypothesis_ (like you're trying to do) No. I am simply saying that there is a lower bound on the rate…

> P=0.05 means that there is a 5% chance that the null is true

No. It simply does not mean that. It would be very convenient if it did, but it does not and can not.

p=0.05 means that if the null hypothesis were true and you ran your experiment, then you would only have a 5% chance of getting the results that you did.

The distinction is not at all obvious, but you'll need to understand it before you can make sense of statistics.

The p-value is not the number you want; what you want is the probability that the null is false, but you can't have that unless you know the actual probability distribution. Which generally isn't something you could possibly know or figure out.

(You can make educated guesses about that distribution, by making a guess and then updating it based on observed results. That's Bayesian statistics, which is great because it does give you a full probability distribution and all you have to feed into it is... uh... a probability distribution.)

Re: The Irreproducibility Crisis of Modern Science

#252
post #68

The site is down so I can't read the original report, but I've read reports on this topic in the past so I'm going to chime in with some "usual suspects" caveats: 1. No result is 100% reproducible because you can never completely reproduce the conditions of any experiment. The best you can hope to do is to reproduce the conditions that matter , but enumerating those has to be part of the theory you are testing, and s…

> The statistical tests currently in widespread use as a criterion for publication in peer-reviewed journals guarantee that at least one result in 20 will be due to chance and not because the hypothesis being tested is actually true. I have a pedantic correction, but a relevant one. If a journal makes a rule that results must have p < .05, and then scientists go off and do a bunch of well-managed science they will ge…

Is it? You're capped at 0.05. You can't print "obvious" results. ("With a confidence of p=0.00002173, we find that people are more likely to be angry in the slapped-in-the-face-by-the-researcher group than the control group...") So there's a natural floor determined by what is worth doing research on.

Now, it turns out that the floor is far lower than researchers believe and the field would be better off by slower incremental progress built atop "obvious" results, but still, I assert that publishability is an adequate explanation for high p values and obscures the evidence for bad practices.

Not that I have any doubt that bad practices abound; I have an undergraduate degree in psychology and am well aware of how things are done. It's not even malicious; people don't understand statistics, don't understand significance, don't understand models, and are mostly just doing whatever they can to make themselves feel good about their own intuitions and biases.

Re: The Irreproducibility Crisis of Modern Science

#253
post #251
post #229

Earlier quoted context omitted.

> it can’t evaluate which of two competing cases is more likely Of course it can. In fact, it can tell you exactly how much more likely one case is versus the other. P=0.05 means that there is a 5% chance that the null is true and hence a 95% chance the the null is false. > says something about the likelihood of the _hypothesis_ (like you're trying to do) No. I am simply saying that there is a lower bound on the rate…

> P=0.05 means that there is a 5% chance that the null is true No. It simply does not mean that. It would be very convenient if it did, but it does not and can not. p=0.05 means that if the null hypothesis were true and you ran your experiment, then you would only have a 5% chance of getting the results that you did. The distinction is not at all obvious, but you'll need to understand it before you can make sense of…

[deleted]

Re: The Irreproducibility Crisis of Modern Science

#254
post #251
post #229

Earlier quoted context omitted.

> it can’t evaluate which of two competing cases is more likely Of course it can. In fact, it can tell you exactly how much more likely one case is versus the other. P=0.05 means that there is a 5% chance that the null is true and hence a 95% chance the the null is false. > says something about the likelihood of the _hypothesis_ (like you're trying to do) No. I am simply saying that there is a lower bound on the rate…

> P=0.05 means that there is a 5% chance that the null is true No. It simply does not mean that. It would be very convenient if it did, but it does not and can not. p=0.05 means that if the null hypothesis were true and you ran your experiment, then you would only have a 5% chance of getting the results that you did. The distinction is not at all obvious, but you'll need to understand it before you can make sense of…

Yes, you're right. I was making the tacit assumption that the null is true most of the time. On this assumption, the two statements are more or less equivalent. I believe this assumption is correct. If it weren't, making real scientific discoveries would be a lot easier. But it is an assumption.

Re: The Irreproducibility Crisis of Modern Science

#255
post #248

Earlier quoted context omitted.

> What does that have to do with reproducibility? That is exactly what I wanted to ask you originally when you cited the celestial mechanics example. It is not relevant in cases we are deriving a theory underlying the behavior , rather than axiomizing the very specific behavior itself...

I don't understand the difference between "deriving a theory underlying the behavior" and "axiomizing [sic] the very specific behavior itself."

Suppose you see an object A, moving straight line getting caught in the gravity of object B, ending up in a orbit around it.

Deriving an underlying theory would be deducing the force of gravity from it.

Axiomizing the very specific behavior would be making a rule that says. "If the object A (and object A only), while moving in a straight line, comes at so and so coordinates with respective to B, will end up in an orbit around it"

Re: The Irreproducibility Crisis of Modern Science

#256
post #248

Earlier quoted context omitted.

I don't understand the difference between "deriving a theory underlying the behavior" and "axiomizing [sic] the very specific behavior itself."

Suppose you see an object A, moving straight line getting caught in the gravity of object B, ending up in a orbit around it. Deriving an underlying theory would be deducing the force of gravity from it. Axiomizing the very specific behavior would be making a rule that says. "If the object A (and object A only), while moving in a straight line, comes at so and so coordinates with respective to B, will end up in an orb…

Ah.

So these are two completely orthogonal issues. The question of what you do with the result of an experiment is completely independent of whether or not you can reproduce that result. The latter is what is under discussion here.

So you can, for example, observe that there was a solar eclipse at a particular time and a particular place. But you cannot repeat that observation. (By way of contrast, you can drop and object from a particular height and observe that it takes a particular time to fall. You can repeat that observation.)

What you do with those observations has (almost) nothing to do with whether or not you can repeat them.

Re: The Irreproducibility Crisis of Modern Science

#257
post #254
post #251

Earlier quoted context omitted.

> P=0.05 means that there is a 5% chance that the null is true No. It simply does not mean that. It would be very convenient if it did, but it does not and can not. p=0.05 means that if the null hypothesis were true and you ran your experiment, then you would only have a 5% chance of getting the results that you did. The distinction is not at all obvious, but you'll need to understand it before you can make sense of…

Yes, you're right. I was making the tacit assumption that the null is true most of the time. On this assumption, the two statements are more or less equivalent. I believe this assumption is correct. If it weren't, making real scientific discoveries would be a lot easier. But it is an assumption.

You're falling prey to the very common "p-value fallacy." P-values talk about the likelihood of the data, not the likelihood of the hypothesis. The two are unfortunately not linked.

Re: The Irreproducibility Crisis of Modern Science

#258
post #212

Earlier quoted context omitted.

> 1. No result is 100% reproducible because you can never completely reproduce the conditions of any experiment. This is why you include the error in your results. We don't care if your two experiments result in 100% the same results. We care about if your results line up within the error of your experiment. > 2. Even a completely non-reproducible result can be scientifically significant. For example, celestial event…

> Flipping 10 heads doesn't guarantee 5 tails I never said it did. But in the long run, flipping a fair coin will generate pretty close to 50% heads and 50% tails. That's what it means to be a fair coin. Likewise, in the long run, using a p threshold of 0.05 (which many journals do) will generate 5% false positive results (that's that the 0.05 means ), i.e. on average 1 false positive result for every 20 experiments…

How can u get a wrong positive result if you only test correct H1s?

Re: The Irreproducibility Crisis of Modern Science

#259
post #212

Earlier quoted context omitted.

> Flipping 10 heads doesn't guarantee 5 tails I never said it did. But in the long run, flipping a fair coin will generate pretty close to 50% heads and 50% tails. That's what it means to be a fair coin. Likewise, in the long run, using a p threshold of 0.05 (which many journals do) will generate 5% false positive results (that's that the 0.05 means ), i.e. on average 1 false positive result for every 20 experiments…

How can u get a wrong positive result if you only test correct H1s?

Because there is a very long causal chain between raw physical phenomena and your perceptions. Your equipment could be faulty. You could be suffering from hallucinations. You could be a brain in a vat.

Example: I want to test if evolution is true. I hypothesize that if evolution is not true, then God will give me some sort of sign, e.g. I pray to God to make a coin that I flip come up heads if evolution is false. I flip the coin and it comes up tails. I conclude that evolution is true. That conclusion would be wrong even if evolution is true.

Re: The Irreproducibility Crisis of Modern Science

#260
post #257
post #254

Earlier quoted context omitted.

Yes, you're right. I was making the tacit assumption that the null is true most of the time. On this assumption, the two statements are more or less equivalent. I believe this assumption is correct. If it weren't, making real scientific discoveries would be a lot easier. But it is an assumption.

You're falling prey to the very common "p-value fallacy." P-values talk about the likelihood of the data, not the likelihood of the hypothesis. The two are unfortunately not linked.

Yes, I am aware of this. P-values tell you the probability that what looks like a positive result is not in fact a positive result but merely a fluke, i.e. the p-value is (a lower bound on) the probability of a false positive. So?
Post reply on HN