Live data from Hacker News

P values are not as reliable as many scientists assume (2014)

nature.com

71–80 of 90 posts

Re: P values are not as reliable as many scientists assume (2014)

#71
post #66

Earlier quoted context omitted.

> But again, when the null-hypothesis doesn't hold, p-value tells you very little (it's actually undefined in the math). The p-value is well defined whether the null hypothesis holds or not. You calculate it assuming it does. There you go, you have a properly calculated p-value. That's what physicists do: "Taking into account the entire mass range of the search, 110– 600 GeV, the global significance of the excess is…

Well, before you go, I implore you to look into the actual computation and theory of 'p-value'. A p-value is simply P(X|H). P(X|H) only means something when H is true. If H is false, P(X|H) tells you nothing. Since H is your null-hypothesis, if it does not actually hold in the real-world, P(X|H) is meaningless. If you read the paper I linked, they never explicitly call out the null hypothesis (nor do, I believe, they…

> A p-value is simply P(X|H). P(X|H) only means something when H is true. If H is false, P(X|H) tells you nothing.

If you know H is false, P(X|H) tells you nothing. But then H wouldn't be a hypothesis, null or otherwise.

If you don't know whether H is true, but you do know something about X, P(X|H) tells you something useful about whether the positive hypothesis to which H is the alternative has an effect apparent in the world to explain.

> The null hypothesis can never be 'rejected' (ie. p-value can never reach 0).

Rejection of the null hypothesis does not mean p-value = 0. Scientific progress is not based on logical certainty, but rather practical utility.

Necessary truths are the domain of pure logic, not empirical science.

Re: P values are not as reliable as many scientists assume (2014)

#72
post #70

Earlier quoted context omitted.

Well, before you go, I implore you to look into the actual computation and theory of 'p-value'. A p-value is simply P(X|H). P(X|H) only means something when H is true. If H is false, P(X|H) tells you nothing. Since H is your null-hypothesis, if it does not actually hold in the real-world, P(X|H) is meaningless. If you read the paper I linked, they never explicitly call out the null hypothesis (nor do, I believe, they…

I think we agree that their null hypothesis is "there is a background, with events coming from all the known particles". I think we agree that their conclusion is "these results provide conclusive evidence for the discovery of a new particle". I don't see how can they say that there is a new particle without rejecting the hypothesis that there is no such new particle. Of course you can say that the null hypothesis ca…

If the null-hypothesis is true, every observation made should be consistent with it. This will result in P(X|H) trending to 1 (since there will be experimental variance). No other experimental design makes sense.

>The meaning [of p-value when H is false] is clear: "the probability of getting a value for the statistic as high as the observed one if the null hypothesis was true"

This is a logical fallacy. It is counterfactual to consider a world where the null hypothesis is true, when it is not.

In fact, this is precisely the feature of the universe that p-value based experimentation exploits and is essentially the only way for us to gain any information about 'reality'.

>By definition, if the null hypothesis is true the p-value is uniformly distributed between 0 and 1.

I don't think so. If that were true, p-value would be entirely useless.

Re: P values are not as reliable as many scientists assume (2014)

#73
post #70

Earlier quoted context omitted.

I think we agree that their null hypothesis is "there is a background, with events coming from all the known particles". I think we agree that their conclusion is "these results provide conclusive evidence for the discovery of a new particle". I don't see how can they say that there is a new particle without rejecting the hypothesis that there is no such new particle. Of course you can say that the null hypothesis ca…

If the null-hypothesis is true, every observation made should be consistent with it. This will result in P(X|H) trending to 1 (since there will be experimental variance). No other experimental design makes sense. >The meaning [of p-value when H is false] is clear: "the probability of getting a value for the statistic as high as the observed one if the null hypothesis was true" This is a logical fallacy. It is counter…

> If the null-hypothesis is true, every observation made should be consistent with it. This will result in P(X|H) trending to 1

Every observation being consistent with H doesn't mean that for each event X that occurs, the conditional probability of X given H will be, or "trend toward", 1. Assuming a perfectly deterministic universe, the P(X|everything else that is true) will be 1 for every X that occurs, but that doesn't mean P(X|H) for any particular true proposition H will be anything like that.

Re: P values are not as reliable as many scientists assume (2014)

#74

Earlier quoted context omitted.

Well, before you go, I implore you to look into the actual computation and theory of 'p-value'. A p-value is simply P(X|H). P(X|H) only means something when H is true. If H is false, P(X|H) tells you nothing. Since H is your null-hypothesis, if it does not actually hold in the real-world, P(X|H) is meaningless. If you read the paper I linked, they never explicitly call out the null hypothesis (nor do, I believe, they…

> A p-value is simply P(X|H). P(X|H) only means something when H is true. If H is false, P(X|H) tells you nothing. If you know H is false, P(X|H) tells you nothing. But then H wouldn't be a hypothesis, null or otherwise. If you don't know whether H is true, but you do know something about X, P(X|H) tells you something useful about whether the positive hypothesis to which H is the alternative has an effect apparent in…

What is the interpretation of a p-value = 0 then? Empirical science can never reject any theory, it is not powerful enough. At best it can provide a selection of least worst explanations.

The H in P(X|H) does not mean 'assumed to be true', it means 'is in fact true'. If H is in fact false, it is counter-factual to assume it is true and therefore any conclusions drawn from the assumption are invalid. This is independent of belief in H.

>Necessary truths are the domain of pure logic, not empirical science.

Science can never deliver truth, which is why it can never truly reject anything (including null-hypotheses).

More generally, this is referred to as the problem of induction.

Re: P values are not as reliable as many scientists assume (2014)

#75

Earlier quoted context omitted.

If the null-hypothesis is true, every observation made should be consistent with it. This will result in P(X|H) trending to 1 (since there will be experimental variance). No other experimental design makes sense. >The meaning [of p-value when H is false] is clear: "the probability of getting a value for the statistic as high as the observed one if the null hypothesis was true" This is a logical fallacy. It is counter…

> If the null-hypothesis is true, every observation made should be consistent with it. This will result in P(X|H) trending to 1 Every observation being consistent with H doesn't mean that for each event X that occurs, the conditional probability of X given H will be, or "trend toward", 1. Assuming a perfectly deterministic universe, the P(X|everything else that is true) will be 1 for every X that occurs, but that doe…

>Every observation being consistent with H doesn't mean that for each event X that occurs, the conditional probability of X given H will be, or "trend toward", 1.

Agreed. If you plot P(X|H) (computed over all observations) over time, in a well-formed experiment the line will trend to 1 if the null-hypothesis is true and to 0 if the null-hypothesis is false.

It really is that simple.

Re: P values are not as reliable as many scientists assume (2014)

#76
post #69
post #68

Earlier quoted context omitted.

First of all, thank you for responding, I expected that no one would. Second of all, you point out that "no false negatives" is not always realistic. Fair enough. That means that 4.8% need not be the right answer. But at least it's clear how I got it. I still have no clue how on earth they got 29%, rather than 28% or 31%. Am I missing something?

1/(1-1/(e p log(p)))) This is formula (3) in the linked paper. I agree that the derivation is not obvious, but probably you should get the same result using the calculation I described for different values of mu and looking for the minimum.

thanks!

Re: P values are not as reliable as many scientists assume (2014)

#77
post #70

Earlier quoted context omitted.

I think we agree that their null hypothesis is "there is a background, with events coming from all the known particles". I think we agree that their conclusion is "these results provide conclusive evidence for the discovery of a new particle". I don't see how can they say that there is a new particle without rejecting the hypothesis that there is no such new particle. Of course you can say that the null hypothesis ca…

If the null-hypothesis is true, every observation made should be consistent with it. This will result in P(X|H) trending to 1 (since there will be experimental variance). No other experimental design makes sense. >The meaning [of p-value when H is false] is clear: "the probability of getting a value for the statistic as high as the observed one if the null hypothesis was true" This is a logical fallacy. It is counter…

You definitely do not know what a p-value is. When you wrote "P(X|H)" I though X was shorthand for T>T(X) where T is the statistic, not that you were referring to the actual data X.

P(X|H) doesn't have the properties you claim, anyway. P(X|H)=1 corresponds to the case where only one outcome is possible. In non-trivial cases, the more data you add the lower this number will be.

Assume H="you have a fair coin". You throw it once: heads. P(X|H)=P(h|faircoin)=1/2 You throw it again: tails. P(X|H)=P(ht|faircoin)=1/4. You throw it again: tails. P(X|H)=P(htt|faircoin)=1/8. I guess the experiment is not well-formed...

Re: P values are not as reliable as many scientists assume (2014)

#78
post #77

Earlier quoted context omitted.

If the null-hypothesis is true, every observation made should be consistent with it. This will result in P(X|H) trending to 1 (since there will be experimental variance). No other experimental design makes sense. >The meaning [of p-value when H is false] is clear: "the probability of getting a value for the statistic as high as the observed one if the null hypothesis was true" This is a logical fallacy. It is counter…

You definitely do not know what a p-value is. When you wrote "P(X|H)" I though X was shorthand for T>T(X) where T is the statistic, not that you were referring to the actual data X. P(X|H) doesn't have the properties you claim, anyway. P(X|H)=1 corresponds to the case where only one outcome is possible. In non-trivial cases, the more data you add the lower this number will be. Assume H="you have a fair coin". You thr…

X can be any random variable that satisfies the requirements of the null-hypothesis.

A more appropriate variable for your experiment would probably be the ratio of heads to tails (may need to add a bias to avoid division by 0).

"you have a fair coin" is not a hypothesis, at least not a well-defined one.

Re: P values are not as reliable as many scientists assume (2014)

#79
post #61
post #55

I'm probably commenting too late to get my question answered, but here goes: the article has a pretty picture where they show how likely your p-values will mislead you depending on how likely the null hypothesis is. For instance, they say if you think that the null hypothesis has a 50% probability of being right, and you get p=5%, then there's still a 29% chance the null hypothesis is true. But according to my calcul…

Assume that you're sampling from a normal distribution with known standard deviation sigma (1 for simplicity) and unknown mean mu. To test if the mean is larger than (the null hypothesis) mu=0 you can check if the observed value is larger than 1.64 sigma (for the 95% confidence test). So if your observation is larger than 1.64 you reject the null hypothesis. Your calculation would be correct only if the assumption "n…

I think I'm still missing something here. In particular, I'm still not getting 29% as a lower bound. I'm getting around 20%. If we compare mu=0 with mu=1.64, the probability density at x=1.64 is roughly 0.1 and 0.4, respectively, so the lower bound should be .1/(.1+.4)=1/5. No? Unless they were assuming something other than "two normal distributions with the same variance"?

Re: P values are not as reliable as many scientists assume (2014)

#80
post #77

Earlier quoted context omitted.

You definitely do not know what a p-value is. When you wrote "P(X|H)" I though X was shorthand for T>T(X) where T is the statistic, not that you were referring to the actual data X. P(X|H) doesn't have the properties you claim, anyway. P(X|H)=1 corresponds to the case where only one outcome is possible. In non-trivial cases, the more data you add the lower this number will be. Assume H="you have a fair coin". You thr…

X can be any random variable that satisfies the requirements of the null-hypothesis. A more appropriate variable for your experiment would probably be the ratio of heads to tails (may need to add a bias to avoid division by 0). "you have a fair coin" is not a hypothesis, at least not a well-defined one.

Ok, so you're thinking about a random variable which converges to some value when the null hypothesis is true. This is fine, but it has nothing to do whatsoever with p-values.

Let me say that your notation is not very appropriate. It makes no sense to say that P(X|H) converges to 1. If you expect X to converge to C if the null hypothesis is true, you can simply say X->C. A proper notation involving probabilities would be P(|X-C|>epsilon)->0 for any positive epsilon (convergence in probability) or maybe P(X->C)=1 (convergence almost surely).

Taking as you suggest X=(#tails/#heads), you expect that X->1 if the coin is fair (I'm not sure why you find this is not a well-defined null hypothesis, but I don't really care). However, P(X)0 as the number if trials increases (X will get closer to 1 on average, but getting exactly 1 will get more and more unlikely).

As I said, you're free to prefer your converging statistics and your well-defined null hypothesis. But you should be aware that people are talking about something completely different when discussing things like the 1e-7 p-value in the Higgs boson discovery or the reproducibility of statistically significant results.

EDIT: Another example, maybe better-defined: a random variable distributed (under the null hypothesis) x~Normal(mu=0,sigma=1). Let's say you take N samples (I let you choose the number, so I don't pick one which is not good enough).The statistic is the mean X=(x_1+x_2+..+x_N)/N. If the null hypothesis is true, X->mu=0. You get X=1/sqrt(N). What's your "p-value" in that case?

Post reply on HN