Live data from Hacker News

P values are not as reliable as many scientists assume (2014)

nature.com

31–40 of 90 posts

Re: P values are not as reliable as many scientists assume (2014)

#31
post #28

Earlier quoted context omitted.

No. P-values don't work that way and don't mean what you think they mean. Read OP or heck, any of the classics like "Why most published research findings are false" http://dx.plos.org/10.1371/journal.pmed.0020124 (36% may or may not be bad, but you can't know without additional stuff like power or prior probability of hypotheses being true; p-values have no intuitive meaning and aren't an answer to any question that…

A proper rebuttal would show what a p-value actually is and how it differs from what I claimed. Now, since a p-value is exactly what I previously claimed, you obviously can't do that. I'm not even sure what you are arguing against me here.

The p-value is the chance of a false positive. But you don't know what the rate of true positives is, or the rate of false negatives.

In a world where there are only false positives and true negatives, and people publish all positive and negative results, then reproduction of a paper should be 95%.

But the reproduction rate when there actually is an effect is not 95%. Depending on sample size, I might get a true positive 20% of the time and a false negative 80% of the time, or I might get a true positive 99.8% of the time and a false negative .2% of the time.

So the average reproduction rate, where an effect actually exists, can be almost any number between 5 and 100. There is no reason to assume it will be 95%.

So the average reproduction rate, where some effects are real and some are imaginary, will almost certainly not be exactly 95%, and that is not a problem in and of itself.

(And when you talk about an average p-value of .05, that sounds like only publishing positive results, which is blatantly going to fail reproduction. 100 false hypotheses -> 5 publications, all false positives -> 5% reproduction rate)

Re: P values are not as reliable as many scientists assume (2014)

#32

Earlier quoted context omitted.

A proper rebuttal would show what a p-value actually is and how it differs from what I claimed. Now, since a p-value is exactly what I previously claimed, you obviously can't do that. I'm not even sure what you are arguing against me here.

The p-value is the chance of a false positive. But you don't know what the rate of true positives is, or the rate of false negatives. In a world where there are only false positives and true negatives, and people publish all positive and negative results, then reproduction of a paper should be 95%. But the reproduction rate when there actually is an effect is not 95%. Depending on sample size, I might get a true posi…

>In a world where there are only false positives and true negatives, and people publish all positive and negative results, then reproduction of a paper should be 95%.

This is the world p-value assumes and is therefore the only one worth considering in relation to my comment.

If an experiment is not well-formed then of course you won't see reproduction at the expected rate. This is what I'm referring to when I say that the low reproduction rate points to deep, fundamental flaws in the experiments.

I agree that the reproduction rate will never be exactly 95% (or 1 - p) due to the discrete nature of experimentation [that's why I used a ~ in front :)], but the reproduction rate of a well-formed experiment should very closely track 1 - p.

Re: P values are not as reliable as many scientists assume (2014)

#33

Earlier quoted context omitted.

The p-value is the chance of a false positive. But you don't know what the rate of true positives is, or the rate of false negatives. In a world where there are only false positives and true negatives, and people publish all positive and negative results, then reproduction of a paper should be 95%. But the reproduction rate when there actually is an effect is not 95%. Depending on sample size, I might get a true posi…

>In a world where there are only false positives and true negatives, and people publish all positive and negative results, then reproduction of a paper should be 95%. This is the world p-value assumes and is therefore the only one worth considering in relation to my comment. If an experiment is not well-formed then of course you won't see reproduction at the expected rate. This is what I'm referring to when I say tha…

>This is the world p-value assumes and is therefore the only one worth considering in relation to my comment.

I'm not sure if that was clear enough. In that world, no one has ever had a hypothesis that was correct. The whole field is useless, measuring things that are wrong and getting the occasional false positive.

You can talk about that world if you want, but it has no connection to reality. It's not p-values that assume that world, it's your misunderstanding of p-values.

>If an experiment is not well-formed then of course you won't see reproduction at the expected rate. This is what I'm referring to when I say that the low reproduction rate points to deep, fundamental flaws in the experiments.

Experiments don't have to have enormous sample sizes to be well-formed. That's the whole point of having a cutoff value.

It's not like an experiment that reproduces 80% of the time disproves the result the rest of the time, it just doesn't quite reach .05 on those trials

>the reproduction rate of a well-formed experiment should very closely track 1 - p

I'm suspicious of this. I don't have time to do the math right now, but an experiment that averages .01 might clear a .05 hurdle far more than 99% of the time, and would definitely be well-formed. And if you set a hurdle at .01 it would only clear it half the time, but it would still be well-formed.

Re: P values are not as reliable as many scientists assume (2014)

#34

Earlier quoted context omitted.

>In a world where there are only false positives and true negatives, and people publish all positive and negative results, then reproduction of a paper should be 95%. This is the world p-value assumes and is therefore the only one worth considering in relation to my comment. If an experiment is not well-formed then of course you won't see reproduction at the expected rate. This is what I'm referring to when I say tha…

>This is the world p-value assumes and is therefore the only one worth considering in relation to my comment. I'm not sure if that was clear enough. In that world, no one has ever had a hypothesis that was correct. The whole field is useless, measuring things that are wrong and getting the occasional false positive. You can talk about that world if you want, but it has no connection to reality. It's not p-values that…

Hypotheses can never be proven to be correct. I don't want to be in any world where it is believed that a hypothesis is or could be correct.

This is a fundamental tenant of science. All that can be done is to reject hypotheses.

You (along with Gwern) have now claimed that I don't understand p-values, but you present no alternative understanding. The reason, of course, is that when you look at the mathematics behind p-value, it is obvious that it is exactly as I claim.

Edit to address your edit:

>I'm suspicious of this. I don't have time to do the math right now, but an experiment that averages .01 might clear a .05 hurdle far more than 99% of the time, and would definitely be well-formed. And if you set a hurdle at .01 it would only clear it half the time, but it would still be well-formed.

You are right that you need to be careful here about what you are comparing across instances. There will be variability since you are only sampling a distribution (most likely at a very low rate) and not observing the entire distribution (which, for continuous distributions, is impossible).

Re: P values are not as reliable as many scientists assume (2014)

#35
post #28

Earlier quoted context omitted.

If the p-values were accurate and averaged around 0.05, ~95% of results should be reproducible. That only 36% were points to deep, fundamental errors.

No. P-values don't work that way and don't mean what you think they mean. Read OP or heck, any of the classics like "Why most published research findings are false" http://dx.plos.org/10.1371/journal.pmed.0020124 (36% may or may not be bad, but you can't know without additional stuff like power or prior probability of hypotheses being true; p-values have no intuitive meaning and aren't an answer to any question that…

I agree, 36% ain't too bad. But,it requires that any literature you use in your research should have been reproduced a few times by other researchers.

Re: P values are not as reliable as many scientists assume (2014)

#36
It is not that P-values are now bad by definition. It's only that they are many times wrongly intepreted. Putting too much confidence in P-values only might result in some wrong conclusions. And this is what some meta analyses discover. Many scientists try hard only to reach the "golden" <0.05 in order to claim discovery and publish it. This is why there is so many papers that misteriously cluster around 0.05...

Re: P values are not as reliable as many scientists assume (2014)

#37
post #6

Earlier quoted context omitted.

[deleted]

First off, there's no need for all-caps. This isn't 4chan. > The fact that assumptions are considered as some unavoidable, forgivable, intricate part of science is part of what fuels anti-science and politics. No one here, as far as I can tell, is saying, 'oh well, science is full of assumptions therefore science is invalid.' The problem is not with science in general being valid or invalid, but rather with the sorts…

>First off, there's no need for all-caps. This isn't 4chan.

Oh man, it looks like I missed something priceless.

Re: P values are not as reliable as many scientists assume (2014)

#38

Earlier quoted context omitted.

>This is the world p-value assumes and is therefore the only one worth considering in relation to my comment. I'm not sure if that was clear enough. In that world, no one has ever had a hypothesis that was correct. The whole field is useless, measuring things that are wrong and getting the occasional false positive. You can talk about that world if you want, but it has no connection to reality. It's not p-values that…

Hypotheses can never be proven to be correct. I don't want to be in any world where it is believed that a hypothesis is or could be correct. This is a fundamental tenant of science. All that can be done is to reject hypotheses. You (along with Gwern) have now claimed that I don't understand p-values, but you present no alternative understanding. The reason, of course, is that when you look at the mathematics behind p…

On a certain philosophical level you can never be absolutely sure of anything, and p-values are meaningless.

On a practical level, p-values are the chance that a correlation is reported where 'reality' does not have a correlation. This is not the same number as the chance that the result agrees with 'reality'.

You can reject the concept of objectivity, but you cannot reject that logic. So I have explained the alternative understanding fine, just go back and replace 'true' and 'false' and 'correct' with a philosophically-hedged version.

Re: P values are not as reliable as many scientists assume (2014)

#39

"Essentially, all models are wrong, but some are useful." --George E.P. Box

The p-value test isn't a model, it's a measure of the significance of an effect in data against random noise.

Which arises from a model (!) of random noise and of your effect.

Re: P values are not as reliable as many scientists assume (2014)

#40

Earlier quoted context omitted.

Hypotheses can never be proven to be correct. I don't want to be in any world where it is believed that a hypothesis is or could be correct. This is a fundamental tenant of science. All that can be done is to reject hypotheses. You (along with Gwern) have now claimed that I don't understand p-values, but you present no alternative understanding. The reason, of course, is that when you look at the mathematics behind p…

On a certain philosophical level you can never be absolutely sure of anything, and p-values are meaningless. On a practical level, p-values are the chance that a correlation is reported where 'reality' does not have a correlation. This is not the same number as the chance that the result agrees with 'reality'. You can reject the concept of objectivity, but you cannot reject that logic. So I have explained the alterna…

On a practical level, people may not be able execute a well formed experiment. I completely agree with that.

However, that doesn't change the meaning of the mathematics, only that your reality has diverged from what you originally intended/believed.

What is the meaning of the number that people call 'p-value' when it is not calculated on a well-formed experiment? I'm not sure if there is a general formula, but you may be able to find some meaning in a particular instance.

Post reply on HN