Live data from Hacker News

P values are not as reliable as many scientists assume (2014)

nature.com

61–70 of 90 posts

Re: P values are not as reliable as many scientists assume (2014)

#61
post #55

I'm probably commenting too late to get my question answered, but here goes: the article has a pretty picture where they show how likely your p-values will mislead you depending on how likely the null hypothesis is. For instance, they say if you think that the null hypothesis has a 50% probability of being right, and you get p=5%, then there's still a 29% chance the null hypothesis is true. But according to my calcul…

Assume that you're sampling from a normal distribution with known standard deviation sigma (1 for simplicity) and unknown mean mu. To test if the mean is larger than (the null hypothesis) mu=0 you can check if the observed value is larger than 1.64 sigma (for the 95% confidence test). So if your observation is larger than 1.64 you reject the null hypothesis.

Your calculation would be correct only if the assumption "no false negatives" is approximately valid. This is the case when the true value is large in terms of sigma (say mu=6). Then for the 100 cases with mu=0 you'll reject the null 5 times on average, and for each one of the 100 cases with mu=100 you will reject the null (unless you're unlucky: there will be a false negative around once in 150000 trials).

But you're conditioning on pWhen the true value of mu gets closer to 0, you cannot ignore the false negatives. For example if mu=0.1 the rejection rate will be quite similar to the mu=0 case (the probability of getting 0.4Somewhere between the two extreme cases, there is a lower bound for this "false discovery rate".

See http://faculty.washington.edu/jonno/SISG-2011/lectures/sellk... and in particular figure 2.

Re: P values are not as reliable as many scientists assume (2014)

#62
post #39

Earlier quoted context omitted.

Which arises from a model (!) of random noise and of your effect.

I see - my mistake. That's a very broad definition of 'model' though isn't it? Including 'random numbers'? You might as well say everything is a model in which case the original quote says nothing :-)

It's perhaps a bit like "everything is a model" in the sense that all of these tests, even the model-free ones, arise from a coherent choice of assumptions and, if you for a moment take the Bayesian perspective very seriously, prior distributions over conditionals. The original quote should be taken to mean that any particular choice of assumptions is limiting, but making interesting choices can drive interesting questions which are thought provoking and meaningful even if they are wrong.

Re: P values are not as reliable as many scientists assume (2014)

#63
post #59

Earlier quoted context omitted.

Measuring a p-value is equivalent to calculating a p-value (ie. calculate the conditional probability P(X|H)). I don't really agree that your die experiment is well-formed. For one, you are grossly under-sampling. It's known a priori that there are at least six possible outcomes, yet you are only considering three rolls, so you don't even have the possibility of observing each distinct value even once. The p-value of…

It's clear that you have your own concept of a p-value, which is quite different from the one used by all the other people (including the proper interpretation and the usual misinterpretations). You disagree with all the provided examples, but you have not given any concrete example of how the p-value would be used in a "well-formed experiment" (another concept that seems unique to you). Of course you're free to rede…

I have redefined nothing. Go pull the the Higgs data; it will be as I say. Go read how to form experiments and calculate p-values, nothing will be substantially different than what I have said here.

Re: P values are not as reliable as many scientists assume (2014)

#64
post #59

Earlier quoted context omitted.

It's clear that you have your own concept of a p-value, which is quite different from the one used by all the other people (including the proper interpretation and the usual misinterpretations). You disagree with all the provided examples, but you have not given any concrete example of how the p-value would be used in a "well-formed experiment" (another concept that seems unique to you). Of course you're free to rede…

I have redefined nothing. Go pull the the Higgs data; it will be as I say. Go read how to form experiments and calculate p-values, nothing will be substantially different than what I have said here.

I already explained a few messages ago that the null hypothesis was indeed that there is no Higgs boson and the peak observed in the LHC data is just noise. They made their analysis and rejected the null hypothesis (p-value less than 0.000001).

"If you correctly measure a p-value of 0.05, that measurement explicitly means that you expect that 95% of your future observations to be consistent with the hypothesis which you used to determine that p-value of 0.05."

Did they correctly measure a p-value of 0.000001? They (and everyone else, apart from maybe you) think that they did.

Do you think they expect 95% of their future observations to be consistent with the null hypothesis (that they used to determine that p-value)?

I would say they were quite confident that the observed peak was not noise, and therefore they expected the signal to be there again if the experiment was to be repeated, rejecting again the null hypothesis. Which is why they announced they had discovered the Higgs boson. But maybe you can convince the Swedish Academy of Science to take Higgs' prize back...

Re: P values are not as reliable as many scientists assume (2014)

#65
post #64

Earlier quoted context omitted.

I have redefined nothing. Go pull the the Higgs data; it will be as I say. Go read how to form experiments and calculate p-values, nothing will be substantially different than what I have said here.

I already explained a few messages ago that the null hypothesis was indeed that there is no Higgs boson and the peak observed in the LHC data is just noise. They made their analysis and rejected the null hypothesis (p-value less than 0.000001). "If you correctly measure a p-value of 0.05, that measurement explicitly means that you expect that 95% of your future observations to be consistent with the hypothesis which…

Here is one paper: http://arxiv.org/pdf/1207.7214v2.pdf

Look at Figures 8 and 9. They show the p-value (at whatever level of data was collected when the paper was written) over the parameter space that is being searched. You can see that the values observed have a clear separation -- most are close to 1 (null-hypothesis holds) with just one significant dip towards 0 (null-hypothesis doesn't hold). If you were to animate this graph with the p-values over time (as more observations are made), you would see the trend towards 0 or 1 much clearer.

The Boson experimenters would expect a (1 - p) reproduction rate for the next observation made (if the null-hypothesis holds). That is, the next observation has a p probability of fitting within the parameters of the null-hypothesis and (1-p) probability that it is inconsistent with the null-hypothesis. Why would they expect that? Because the math involved in telling you whether or not that is what you should expect is exactly what p-value calculates (again, assuming a well-formed experiement -- which the Higgs experiments probably are).

But again, when the null-hypothesis doesn't hold, p-value tells you very little (it's actually undefined in the math).

Re: P values are not as reliable as many scientists assume (2014)

#66
post #64

Earlier quoted context omitted.

I already explained a few messages ago that the null hypothesis was indeed that there is no Higgs boson and the peak observed in the LHC data is just noise. They made their analysis and rejected the null hypothesis (p-value less than 0.000001). "If you correctly measure a p-value of 0.05, that measurement explicitly means that you expect that 95% of your future observations to be consistent with the hypothesis which…

Here is one paper: http://arxiv.org/pdf/1207.7214v2.pdf Look at Figures 8 and 9. They show the p-value (at whatever level of data was collected when the paper was written) over the parameter space that is being searched. You can see that the values observed have a clear separation -- most are close to 1 (null-hypothesis holds) with just one significant dip towards 0 (null-hypothesis doesn't hold). If you were to anim…

> But again, when the null-hypothesis doesn't hold, p-value tells you very little (it's actually undefined in the math).

The p-value is well defined whether the null hypothesis holds or not. You calculate it assuming it does. There you go, you have a properly calculated p-value. That's what physicists do:

"Taking into account the entire mass range of the search, 110– 600 GeV, the global significance of the excess is 5.1 σ, which corresponds to p0 = 1.7 × 10−7."

You see, they have calculated a p-value. Does the null hypothesis hold? I don't think they had any expectations consistent with the null hypothesis being true before the experiment. After the experiment they clearly think that the null hypothesis is false:

"These results provide conclusive evidence for the discovery of a new particle with mass 126.0 ± 0.4 (stat) ± 0.4 (sys) GeV."

They don't see any problem in stating a p-value and rejecting the null hypothesis at the same time (in fact, it's because the p-value that they calculated is very small that they conclude that the null hypothesis doesn't hold). Apparently you see a problem, because if the Higgs boson exists and produces the signal in the experiment then the null hypothesis is false and all the p-value calculations they did to get to that conclusion are "wrong".

Anyway, I have no need to convince you of anything. I can live with people being wrong on the internet.

Re: P values are not as reliable as many scientists assume (2014)

#67
post #66

Earlier quoted context omitted.

Here is one paper: http://arxiv.org/pdf/1207.7214v2.pdf Look at Figures 8 and 9. They show the p-value (at whatever level of data was collected when the paper was written) over the parameter space that is being searched. You can see that the values observed have a clear separation -- most are close to 1 (null-hypothesis holds) with just one significant dip towards 0 (null-hypothesis doesn't hold). If you were to anim…

> But again, when the null-hypothesis doesn't hold, p-value tells you very little (it's actually undefined in the math). The p-value is well defined whether the null hypothesis holds or not. You calculate it assuming it does. There you go, you have a properly calculated p-value. That's what physicists do: "Taking into account the entire mass range of the search, 110– 600 GeV, the global significance of the excess is…

Well, before you go, I implore you to look into the actual computation and theory of 'p-value'.

A p-value is simply P(X|H). P(X|H) only means something when H is true. If H is false, P(X|H) tells you nothing. Since H is your null-hypothesis, if it does not actually hold in the real-world, P(X|H) is meaningless.

If you read the paper I linked, they never explicitly call out the null hypothesis (nor do, I believe, they show the work for their calculations). There should be another paper somewhere that describes exactly what it is, in the terms I am using. So, phrases like, "[t]hey don't see any problem in stating a p-value and rejecting the null hypothesis" make me think you have no idea what you're talking about.

The null hypothesis can never be 'rejected' (ie. p-value can never reach 0). I don't think you will find anyone working on the Higgs boson that will claim otherwise.

Re: P values are not as reliable as many scientists assume (2014)

#68
post #61
post #55

I'm probably commenting too late to get my question answered, but here goes: the article has a pretty picture where they show how likely your p-values will mislead you depending on how likely the null hypothesis is. For instance, they say if you think that the null hypothesis has a 50% probability of being right, and you get p=5%, then there's still a 29% chance the null hypothesis is true. But according to my calcul…

Assume that you're sampling from a normal distribution with known standard deviation sigma (1 for simplicity) and unknown mean mu. To test if the mean is larger than (the null hypothesis) mu=0 you can check if the observed value is larger than 1.64 sigma (for the 95% confidence test). So if your observation is larger than 1.64 you reject the null hypothesis. Your calculation would be correct only if the assumption "n…

First of all, thank you for responding, I expected that no one would. Second of all, you point out that "no false negatives" is not always realistic. Fair enough. That means that 4.8% need not be the right answer. But at least it's clear how I got it. I still have no clue how on earth they got 29%, rather than 28% or 31%. Am I missing something?

Re: P values are not as reliable as many scientists assume (2014)

#69
post #68
post #61

Earlier quoted context omitted.

Assume that you're sampling from a normal distribution with known standard deviation sigma (1 for simplicity) and unknown mean mu. To test if the mean is larger than (the null hypothesis) mu=0 you can check if the observed value is larger than 1.64 sigma (for the 95% confidence test). So if your observation is larger than 1.64 you reject the null hypothesis. Your calculation would be correct only if the assumption "n…

First of all, thank you for responding, I expected that no one would. Second of all, you point out that "no false negatives" is not always realistic. Fair enough. That means that 4.8% need not be the right answer. But at least it's clear how I got it. I still have no clue how on earth they got 29%, rather than 28% or 31%. Am I missing something?

1/(1-1/(e p log(p))))

This is formula (3) in the linked paper. I agree that the derivation is not obvious, but probably you should get the same result using the calculation I described for different values of mu and looking for the minimum.

Re: P values are not as reliable as many scientists assume (2014)

#70
post #66

Earlier quoted context omitted.

> But again, when the null-hypothesis doesn't hold, p-value tells you very little (it's actually undefined in the math). The p-value is well defined whether the null hypothesis holds or not. You calculate it assuming it does. There you go, you have a properly calculated p-value. That's what physicists do: "Taking into account the entire mass range of the search, 110– 600 GeV, the global significance of the excess is…

Well, before you go, I implore you to look into the actual computation and theory of 'p-value'. A p-value is simply P(X|H). P(X|H) only means something when H is true. If H is false, P(X|H) tells you nothing. Since H is your null-hypothesis, if it does not actually hold in the real-world, P(X|H) is meaningless. If you read the paper I linked, they never explicitly call out the null hypothesis (nor do, I believe, they…

I think we agree that their null hypothesis is "there is a background, with events coming from all the known particles". I think we agree that their conclusion is "these results provide conclusive evidence for the discovery of a new particle". I don't see how can they say that there is a new particle without rejecting the hypothesis that there is no such new particle. Of course you can say that the null hypothesis can never be rejected (relevant Dilbert strip: http://dilbert.com/strip/2001-10-25) but then they can never discover a new particle either.

Regarding p-values in general, your definition is the same I've been using all along. But I don't think it is meaningless when the null hypothesis does not hold. The meaning is clear: "the probability of getting a value for the statistic as high as the observed one if the null hypothesis was true". For example, there would be one chance in several millions of observing the kind of data they found at the LHC if the Higgs boson didn't exist.

You might want to look into the theory yourself, because the notion of p-values trending towards 1 if the null hypothesis is true is nonsense. By definition, if the null hypothesis is true the p-value is uniformly distributed between 0 and 1. If you have at some point a p-value close to one (or to any other number for that matter) and keep adding data, in the long run it will still be uniformly distributed between 0 and 1.

Post reply on HN