Live data from Hacker News

P values are not as reliable as many scientists assume (2014)

nature.com

51–60 of 90 posts

Re: P values are not as reliable as many scientists assume (2014)

#51
post #48

Earlier quoted context omitted.

95% of the repeated observations that you make (in the same manner as the observations used to calculate a valid p-value of 0.05) will be consistent with the relevant null hypothesis. What other meaning could there be? The result of an experiment is not a p-value, but a series of observations. Those are what need to be compared.

I guess the bit "results should be reproducible" made us think that you were talking about reproducing the previous results (i.e. if the null hypothesis was rejected in the first trial, obtaining again a rejection if the trial was repeated). If I understand your point, you're saying: "If the null hypothesis is true then with 95% probability it won't be rejected. And, independently of the result of the first trial, if…

It's highly relevant to the discussion (interestingly titled: "P-value not as reliable as many scientists assume"), which is entirely about p-value, as I am describing to you the limits of p-value analysis.

That some perform calculations that are not p-value and call them p-value is not exactly my problem to solve. That others perform meta-analyses with numbers that others call p-values, but which aren't actually p-values isn't really my problem either.

I'll say it again. If you correctly measure (exercise left to reader) a p-value of 0.05, that measurement explicitly means that you expect that 95% of your future observations to be consistent with the hypothesis which you used to determine that p-value of 0.05.

Making future observations that are consistent with a known hypothesis is exactly what reproducibility refers to within the context of science.

If you expect 95% of observations (p = 0.05) to be consistent with previous findings, but only 36% are...you did not calculate a valid p-value (or are now testing something other than your hypothesis).

Re: P values are not as reliable as many scientists assume (2014)

#52
post #48

Earlier quoted context omitted.

I guess the bit "results should be reproducible" made us think that you were talking about reproducing the previous results (i.e. if the null hypothesis was rejected in the first trial, obtaining again a rejection if the trial was repeated). If I understand your point, you're saying: "If the null hypothesis is true then with 95% probability it won't be rejected. And, independently of the result of the first trial, if…

It's highly relevant to the discussion (interestingly titled: "P-value not as reliable as many scientists assume"), which is entirely about p-value, as I am describing to you the limits of p-value analysis. That some perform calculations that are not p-value and call them p-value is not exactly my problem to solve. That others perform meta-analyses with numbers that others call p-values, but which aren't actually p-v…

> I'll say it again. If you correctly measure (exercise left to reader) a p-value of 0.05, that measurement explicitly means that you expect that 95% of your future observations to be consistent with the hypothesis which you used to determine that p-value of 0.05.

What do you mean with "measure a p-value"? You make your observation, calculate a statistic (a function of the observation), and look at the distribution of that statistic under the null. The p-value is, by definition, the percentile of the value you got in that distribution (which might or might not be the actual distribution).

You want to check if a die is loaded to yield 6 more often than it should. The null hypothesis is that the die is fair. You can calculate the distribution for the number of 6's in 3 rolls (0: 58%, 1:35%, 2: 7%, 3: 0.5%). You roll the die three times, you get three 6's. The p-value is 0.005. Do you agree? The p-value is 0.005 whether the die is fair (the null hypothesis is true) or loaded. Do you agree?

> Making future observations that are consistent with a known hypothesis is exactly what reproducibility refers to within the context of science.

Scientific experiments are usually about rejecting the null hypothesis. For example, the null hypothesis might be that there is no Higgs boson and the peak observed in the LHC data is just noise. They made their analysis and rejected the null hypothesis (p-value less than 0.000001, do you think they calculated it properly?). In this context, reproducibility means "finding the Higgs boson again if the experiment is repeated" and not "repeat the experiment and get a result consistent with the null hypothesis".

According to your description of the limits of p-value analysis, the only conclusion that physicists should get out of the experiment is that if they do it again they should expect to get results consistent with the null hypothesis (i.e. no Higgs boson) with 95% probability. But they see it as evidence that the null hypothesis is false and the Higgs boson real.

Re: P values are not as reliable as many scientists assume (2014)

#53

Earlier quoted context omitted.

A well formed experiment tests only a null hypothesis. p-value is exactly the probability that you observed X given that the previously stated null hypothesis was true at the time of observation. The value (1 - p-value) is exactly the probability that you will make an observation consistent with your hypothesis (ie. expected replication rate). Wikipedia has a decent treatment that might help: https://en.wikipedia.org…

But the importance of a p-value is showing when it's not the null hypothesis. The only time you get 95% reproduction is a result that says the null hypothesis is true. You're entirely right about that specific case. But this only happens when nothing correlates. (And almost no science has been done, because most things in fact don't correlate.) A result that disagrees with the null hypothesis at .05 does not imply an…

>You're entirely right about that specific case.

In fact, this is the only case that matters. All other (valid) cases can be reduced to a single, null hypothesis design.

p-value is undefined for hypotheses that are not a null hypothesis. It is also undefined for hypotheses which do not hold.

Sure, you can walk through the motions, put some numbers together, and eventually produce a number between 0 and 1. However that does not mean you have computed a p-value. If you are testing a non-null hypothesis you have not computed a p-value. If you are testing a null hypothesis that doesn't hold, you have not computed a p-value.

Re: P values are not as reliable as many scientists assume (2014)

#54

Earlier quoted context omitted.

But the importance of a p-value is showing when it's not the null hypothesis. The only time you get 95% reproduction is a result that says the null hypothesis is true. You're entirely right about that specific case. But this only happens when nothing correlates. (And almost no science has been done, because most things in fact don't correlate.) A result that disagrees with the null hypothesis at .05 does not imply an…

>You're entirely right about that specific case. In fact, this is the only case that matters. All other (valid) cases can be reduced to a single, null hypothesis design. p-value is undefined for hypotheses that are not a null hypothesis. It is also undefined for hypotheses which do not hold. Sure, you can walk through the motions, put some numbers together, and eventually produce a number between 0 and 1. However tha…

The null hypothesis is where nothing happens. You're supposed to be showing evidence against it. If you redefine things so your "null hypothesis" is where something happens, and you're showing evidence for it, you have done something very very wrong, and you should not be using a .05 threshold either.

Re: P values are not as reliable as many scientists assume (2014)

#55
I'm probably commenting too late to get my question answered, but here goes: the article has a pretty picture where they show how likely your p-values will mislead you depending on how likely the null hypothesis is. For instance, they say if you think that the null hypothesis has a 50% probability of being right, and you get p=5%, then there's still a 29% chance the null hypothesis is true. But according to my calculations, the right number should be 1/21 = 4.8%. What am I missing here? Or are they wrong? My calculations are below:

Curious George has 200 fascinating phenomena he wishes to investigate. In reality, 100 of those are real, and the other hundred are mere coincidences. The experiments for the 100 real phenomena all show that "yes, this is for real". (I'm assuming no false negatives.) Most of the 100 experiments that test bogus phenomena show that "this is bogus", but 5 of them achieve a significance of p=5%, as expected. George then runs of to tell the Man in the Yellow Hat about his 105 amazing discoveries. If Yellow Hat Man knows that half of the phenomena that capture George's attention are bogus, he knows that 5/105 = 1/21 = 4.8% of George's discoveries are likely bogus, even though he doesn't know which ones.

Re: P values are not as reliable as many scientists assume (2014)

#56

Earlier quoted context omitted.

A proper rebuttal would show what a p-value actually is and how it differs from what I claimed. Now, since a p-value is exactly what I previously claimed, you obviously can't do that. I'm not even sure what you are arguing against me here.

The p-value is the chance of a false positive. But you don't know what the rate of true positives is, or the rate of false negatives. In a world where there are only false positives and true negatives, and people publish all positive and negative results, then reproduction of a paper should be 95%. But the reproduction rate when there actually is an effect is not 95%. Depending on sample size, I might get a true posi…

"The p-value is the chance of a false positive."

Nope, It's the chance of getting a result as this or more extreme under the assumption of the null hypotheses.

Re: P values are not as reliable as many scientists assume (2014)

#57
post #50
post #22

Earlier quoted context omitted.

A better word for that would be "premise" or "axiom". Premises and axioms are objective exceptions with objective merits. Assumptions are too personal because they infer belief which is purely subjective. Nature doesn't care about what anyone believes and science should never be a democracy. Assumptions also imply some independent existential entity as valid and are self-validating, whereas premises and axioms are hi…

You're making a distinction that I've never heard anyone make before, and I don't think you're making a convincing argument for it now.

The distinction between axiom and assumption is quite clear, so I'll assume you are referring to premise.

The distinction already exists, which is the beauty of words. I am not making this distinction up. I am merely enforcing them as we do as speakers by the words we choose, based on the accuracy of our expressions.

From Popper:

> A theoretical system may be said to be axiomatized if ... (d) necessary, for the same purpose; which means that they should contain no superfluous assumptions [0].

So either Popper is wrong, or theories should not include unnecessary assumptions. But how do we know if an assumption is necessary without doing the science? And after we do it, are we still going to call necessary assumptions assumptions, even with its subjective implications? Only a person is capable of assuming. A theory with assumptions is still subjective.

A premise on the other hand is objective and specific. It's "a proposition supporting or helping to support a conclusion" [1]. It's a simple device in logic that asserts a dependency. Or in other words, a "necessary assumption".

So then would it not be safer to say axiomatized theories have premises, not assumptions? And in Popper's words, yet to be axiomatized theories have assumptions. That is what makes them hypotheses. And so all the words and their distinctions fall cleanly into place (I did not make anything up).

The original article in Nature was written on the premise that science is based on assumptions and that scientists doing the science rely on assumptions. This is not the premise of science, and is incorrect. Refining our word selection to reflect this understanding would be of great service, particularly to the students.

-- [0] https://books.google.com/books?id=cAKCAgAAQBAJ&pg=PA51&lpg=P... [1] Just from the dictionary, as to avoid my own words. http://dictionary.reference.com/browse/premise

Re: P values are not as reliable as many scientists assume (2014)

#58
post #52

Earlier quoted context omitted.

It's highly relevant to the discussion (interestingly titled: "P-value not as reliable as many scientists assume"), which is entirely about p-value, as I am describing to you the limits of p-value analysis. That some perform calculations that are not p-value and call them p-value is not exactly my problem to solve. That others perform meta-analyses with numbers that others call p-values, but which aren't actually p-v…

> I'll say it again. If you correctly measure (exercise left to reader) a p-value of 0.05, that measurement explicitly means that you expect that 95% of your future observations to be consistent with the hypothesis which you used to determine that p-value of 0.05. What do you mean with "measure a p-value"? You make your observation, calculate a statistic (a function of the observation), and look at the distribution o…

Measuring a p-value is equivalent to calculating a p-value (ie. calculate the conditional probability P(X|H)).

I don't really agree that your die experiment is well-formed. For one, you are grossly under-sampling. It's known a priori that there are at least six possible outcomes, yet you are only considering three rolls, so you don't even have the possibility of observing each distinct value even once.

The p-value of a well-formed experiment should converge towards a fixed value as more observations are made. You will experience variance in the computed value due to the inherently discrete nature of experimentation. This will be especially pronounced for the first observations that are made.

I do not know if the Higgs boson experiment is well-formed. If it is well-formed and their null-hypothesis is true, their p-values will trend towards 1.

If their null-hypothesis is not true then the p-values do not mean much and will trend towards 0.

>In this context, reproducibility means "finding the Higgs boson again if the experiment is repeated"

"The Higgs boson exists" is not a valid hypothesis. Usually the null-hypothesis is "the explanation is measurement/background noise". Since that is really the only valid null-hypothesis, it is most likely what they are using.

Re: P values are not as reliable as many scientists assume (2014)

#59
post #52

Earlier quoted context omitted.

> I'll say it again. If you correctly measure (exercise left to reader) a p-value of 0.05, that measurement explicitly means that you expect that 95% of your future observations to be consistent with the hypothesis which you used to determine that p-value of 0.05. What do you mean with "measure a p-value"? You make your observation, calculate a statistic (a function of the observation), and look at the distribution o…

Measuring a p-value is equivalent to calculating a p-value (ie. calculate the conditional probability P(X|H)). I don't really agree that your die experiment is well-formed. For one, you are grossly under-sampling. It's known a priori that there are at least six possible outcomes, yet you are only considering three rolls, so you don't even have the possibility of observing each distinct value even once. The p-value of…

It's clear that you have your own concept of a p-value, which is quite different from the one used by all the other people (including the proper interpretation and the usual misinterpretations).

You disagree with all the provided examples, but you have not given any concrete example of how the p-value would be used in a "well-formed experiment" (another concept that seems unique to you).

Of course you're free to redefine concepts as you please, if it makes you happy or it is useful to you in any other way.

Re: P values are not as reliable as many scientists assume (2014)

#60
post #39

Earlier quoted context omitted.

The p-value test isn't a model, it's a measure of the significance of an effect in data against random noise.

Which arises from a model (!) of random noise and of your effect.

I see - my mistake. That's a very broad definition of 'model' though isn't it? Including 'random numbers'? You might as well say everything is a model in which case the original quote says nothing :-)
Post reply on HN