Live data from Hacker News

Moving to a World Beyond "p < 0.05" (2019)

tandfonline.com

101–110 of 111 posts

Re: Moving to a World Beyond "p < 0.05" (2019)

#101
post #22

Earlier quoted context omitted.

Watching from the sidelines, I’ve always wondered why everything in the life sciences seems to assume unimodal distributions (that is, typically a normal bell curve). Multimodal distributions are everywhere, and we are losing key insights by ignoring this. A classic example is the difference in response between men and women to a novel pharmaceutical. It’s certainly not the case that scientists are not aware of this…

It’s because statistical tests are based on the distribution of the statistic, not the data itself. If the central limit holds, this distribution will be a bell curve as you say

Aye, but there are cases where it doesn't hold. Lognormal and power law distributions are awfully similar in samples, but matters on the margin.

For example, checking account balances are far from a normal distribution!

Re: Moving to a World Beyond "p < 0.05" (2019)

#102
post #97

Earlier quoted context omitted.

bioinformatician here. nobody has intuition or domain knowledge on all ~20,000 protein coding genes in the human body. That's just not a thing. Routinely comparing what a treatment does we do actually get 20,000 p-values. Feed that into FDR correction, filter for p then you see proof borne out that the p-value thresholding business WORKS.

> Feed that into FDR correction, filter for p That's the domain knowledge. p-values are useful not the fixed cut-off. You know that in your research field p < 0.01 has importance.

> You know that in your research field p A p-value does not measure "importance" (or relevance), and its meaning is not dependent on the research field or domain knowledge: it mostly just depends on effect size and number of replicates (and, in this case, due to the need to apply multiple comparison correction for effective FDR control, it depends on the number of things you are testing).

If you take any fixed effect size (no matter how small/non-important or large/important, as long as it is nonzero), you can make the p-value be arbitrarily small by just taking a sufficiently high number of samples (i.e., replicates). Thus, the p-value does not measure effect importance, it (roughly) measures whether you have enough information to be able to confidently claim that the effect is not exactly zero.

Example: you have a drug that reduces people's body weight by 0.00001% (clearly, an irrelevant/non-important effect, according to my domain knowledge of "people's expectations when they take a weight loss drug"); still, if you collect enough samples (i.e., take the weight of enough people who took the drug and of people who took a placebo, before and after), you can get a p-value as low as you want (0.05, 0.01, 0.001, etc.), mathematically speaking (i.e., as long as you can take an arbitrarily high number of samples). Thus, the p-value clearly can't be measuring the importance of the effect, if you can make it arbitrarily low by just having more measurements (assuming a fixed effect size/importance).

What is research field (or domain knowledge) dependent is the "relevance" of the effect (i.e., the effect size), which is what people should be focusing on anyway ("how big is the effect and how certain am I about its scale?"), rather than p-values (a statement about a hypothetical universe in which we assume the null to be true).

Re: Moving to a World Beyond "p < 0.05" (2019)

#103

Earlier quoted context omitted.

Yeah but without a hardline how would you decide what to publish?

Not publishing results with p >= 0.05 is the reason p-values aren't that useful. This is how you get the replication crisis in psychology. The p-value cutoff of 0.05 just means "an effect this large, or larger, should happen by chance 1 time out of 20". So if 19 failed experiments don't publish and the 1 successful one does, all you've got are spurious results. But you have no way to know that, because you don't see…

> "an effect this large, or larger, should happen by chance 1 time out of 20"

More like "an effect this large, or larger, should happen by chance 1 time out of 20 in the hypothetical universe where we already know that the true size of the effect is zero".

Part of the problem of p-values is that most people can't even parse what it means (not saying it's your case). P-values are never a statement about probabilities in the real world, but always a statement about probabilities in a hypothetical world where we all effects are zero.

"Effect sizes", on the other hand, are more directly meaningful and more likely to be correctly interpreted by people on general, particularly if they have the relevant domain knowledge.

(Otherwise, I 100% agree with the rest of your comment.)

Re: Moving to a World Beyond "p < 0.05" (2019)

#104

Earlier quoted context omitted.

Both of those statements are false. Everything has a result. And the p-value is very literally a quantified measure of how interesting a result was. That's the only thing it purports to measure. "Woman gives birth to fish" is interesting because it has a p-value of zero: under the null hypothesis ("no supernatural effects"), a woman can never give birth to a fish.

I ate cheese yesterday and a celebrity died today: P >> 0.05. There is no result and you can't say anything about whether my cheese eating causes or prevents celebrity deaths. You confuse hypothesis testing with P-values.

[deleted]

Re: Moving to a World Beyond "p < 0.05" (2019)

#105

Earlier quoted context omitted.

Both of those statements are false. Everything has a result. And the p-value is very literally a quantified measure of how interesting a result was. That's the only thing it purports to measure. "Woman gives birth to fish" is interesting because it has a p-value of zero: under the null hypothesis ("no supernatural effects"), a woman can never give birth to a fish.

I ate cheese yesterday and a celebrity died today: P >> 0.05. There is no result and you can't say anything about whether my cheese eating causes or prevents celebrity deaths. You confuse hypothesis testing with P-values.

The result is "a celebrity died today". This result is uninteresting because, according to you, celebrities die much more often than one per twenty days.

I suggest reading your comments before you post them.

Re: Moving to a World Beyond "p < 0.05" (2019)

#107

Earlier quoted context omitted.

I don't disagree with this, but this is very not consistent with "ethnicity is not biologically realized", which suffers from the same logical error but in the other direction. I often wonder how many entrenched culture battles could be ~resolved (at least objectively) by fixing people's cognitive variable types.

In the spirit of randomization and simulation, every culture war debate should be repeated at least 200 times, each with randomly assigned definitions of “justice” and “freedom” drawn from an introductory philosophy textbook. Eating meat is wrong, p = 12/200.

Some day it may make for great training data.

Re: Moving to a World Beyond "p < 0.05" (2019)

#108
post #97

Earlier quoted context omitted.

> Feed that into FDR correction, filter for p That's the domain knowledge. p-values are useful not the fixed cut-off. You know that in your research field p < 0.01 has importance.

> You know that in your research field p A p-value does not measure "importance" (or relevance), and its meaning is not dependent on the research field or domain knowledge: it mostly just depends on effect size and number of replicates (and, in this case, due to the need to apply multiple comparison correction for effective FDR control, it depends on the number of things you are testing). If you take any fixed effect…

I get that in general. I was replying to the person taking about bioinformatics and p value as filter.

Re: Moving to a World Beyond "p < 0.05" (2019)

#109
post #85

Earlier quoted context omitted.

Publishing only significant results is a terrible idea in the first place. Publishing should be based on how interesting the design of the experiment was, not how interesting the result was.

P-value doesn't measure interestingness. If p>0.05 there was no result at all.

p-value doesn't measure interestingness directly of course, but I think people generally find nonsignificant results uninteresting because they think the result is not difficult to explain by the definitionally-uninteresting "null hypothesis".

My point was basically that the reputation / carrer / etc of the experimenter should be mostly independent of the study results. Otherwise you get bad incentives. Obviously we have limited ability to do this in practice, but at least we could fix the way journals decide what to publish.

Re: Moving to a World Beyond "p < 0.05" (2019)

#110
post #40

I think the main issue with "p < 0.05" can be summed up as people use it to "prove" a phenomenon exists rather than a pragmatic cutoff to screen for interesting things to investigate further.

Yeah but without a hardline how would you decide what to publish?

Preregistered studies don't have any p values before they're accepted.
Post reply on HN