Live data from Hacker News

Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

fivethirtyeight.com

31–40 of 130 posts

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#31

5% (1 in 20) is a pretty weak threshold to pass. Let's go 5 sigma (p < 3e-7) for discoveries and reserve 0.05 < p < 3e-7 for stuff we should take closer looks at.

That's fine, but then your power is low, and nearly every result you get will be an exaggeration. The more stringent your p value threshold, the more dramatic your results must be to be significant; if your sample size isn't adequate, you'll only get significance if you overestimate the effect. This is an enormously common problem even with current p value thresholds. It's part of the reason why you see dramatic "A c…

You can either have more dramatic results with more stringent p-value OR have less variance, aka more sampling.

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#32
post #8

Is the p-value really not the probability of your results being due to chance? Is that not a perfectly valid definition of it? I suppose 'chance' is a little hand-wavy, but isn't a p-value just the probability of your data given that your hypothesis is false? Isn't that literally and precisely the probability that they occurred by chance?

[deleted]

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#33

For context, I have a degree in statistics, and I did research with Andrew Gelman (one of the statisticians quoted in the article). Glad to see this is gaining traction! I've been saying this for years: the world would actually be in a better place if we just abandoned p-values altogether. Hypothesis testing is taught in introductory statistics courses because the calculations involved are deceptively easy, whereas m…

This paradox is a great example of why I'm uncomfortable with the idea of coming up with a p-value when you're comparing a sample mean to a point value in the first place.

If your null hypothesis amounts to little more than a scalar value that exists in a vacuum, you haven't really got a null hypothesis. If your null hypothesis is that your results will match what was seen in some other data set collected by some other person who may or may not have been using the same protocol as you, your null hypothesis describes a parallel universe and there's no way to draw an apples-to-apples comparison to it and your alternative hypothesis, which concerns data that did not come from that parallel universe.

So I guess my answer to all questions would be, "Mu."

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#34

For context, I have a degree in statistics, and I did research with Andrew Gelman (one of the statisticians quoted in the article). Glad to see this is gaining traction! I've been saying this for years: the world would actually be in a better place if we just abandoned p-values altogether. Hypothesis testing is taught in introductory statistics courses because the calculations involved are deceptively easy, whereas m…

If you're doing a measurement, why have a null hypothesis? Alice should sample the population at random, take the height measurements, calculate the average, plot the distribution, calculate the variance. If the distribution is not sufficiently smooth then continue to take measurements until it's smooth, or unchanging. Then Alice is done discovering all there is to know about the height distribution of the population…

> Alice should sample the population at random, take the height measurements, calculate the average, plot the distribution, calculate the variance

A student-t test is basically a lossy encoding of that information you describe (the sample mean, the variance, and the sample size).

> If the distribution is not sufficiently smooth then continue to take measurements until it's smooth, or unchanging.

From a frequentist standpoint, this would be considered sloppy and bad methodology. You don't keep sampling until you get the results you want (or until the null hypothesis is invalidated).

That said, as I mention in a comment above, the fact that you can keep sampling until the null hypothesis is invalidated (and that it is mathematically always guaranteed to happen eventually) is a big problem with the concept of hypothesis testing in the first place.

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#35
post #21

If they held the same meeting 20 times, would they reach the same conclusion in 19 of those meetings? On a more serious note, I think that the use of the word "significant" to mean "the effect is reasonably likely to exist by some standard" should be abolished. Webster's 1913 dictionary says: > Deserving to be considered; important; momentous; as, a significant event. Statisticians don't use "significant" to mean imp…

This is extremely important and I'm not surprised the media is either ignorant, doesn't understand, or 'stretches the truth' to adhere to their worldview. Statistical significance has NOTHING to do with magnitude of difference!

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#36

For context, I have a degree in statistics, and I did research with Andrew Gelman (one of the statisticians quoted in the article). Glad to see this is gaining traction! I've been saying this for years: the world would actually be in a better place if we just abandoned p-values altogether. Hypothesis testing is taught in introductory statistics courses because the calculations involved are deceptively easy, whereas m…

Thanks! This is one of the best explanations of the problem with using p-values.

Furthermore, I can recall many times when my stats professors would make it very clear what the "right" answer was, despite it being counter-intuitive but without explaining

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#37

The article submitted here leads to the American Statistical Association statement on the meaning of p values,[1] the first such methodological statement ever formally issued by the association. It's free to read and download. The statement summarizes into these main points, with further explanation in the text of the statement. "What is a p-value? "Informally, a p-value is the probability under a specified statistic…

> "3. Scientific conclusions and business or policy decisions should not be based only on whether a p-value passes a specific threshold.

There is no lower threshold at which the data becomes non-predictive?

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#38
post #37

The article submitted here leads to the American Statistical Association statement on the meaning of p values,[1] the first such methodological statement ever formally issued by the association. It's free to read and download. The statement summarizes into these main points, with further explanation in the text of the statement. "What is a p-value? "Informally, a p-value is the probability under a specified statistic…

> "3. Scientific conclusions and business or policy decisions should not be based only on whether a p-value passes a specific threshold. There is no lower threshold at which the data becomes non-predictive?

More likely they mean the opposite: that a statistically significant P value (by whatever threshold you decide to use) should not be used, by itself, to drive policy decisions. Internally, the effect size still matters. Externally, there are numerous other factors that should drive decisionmaking.

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#39

p-value analysis has its big caveats like multiple comparisons, but Bayesian has its own, such as it's extremely hard to calculate priors. Both are challenging to use in difficult analysis and both can be abused.

> such as it's extremely hard to calculate priors.

I think you mean it is difficult to formulate priors? Typically, calculating a prior only involves sampling from a distribution.

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#40

For context, I have a degree in statistics, and I did research with Andrew Gelman (one of the statisticians quoted in the article). Glad to see this is gaining traction! I've been saying this for years: the world would actually be in a better place if we just abandoned p-values altogether. Hypothesis testing is taught in introductory statistics courses because the calculations involved are deceptively easy, whereas m…

If Alice just wants to know the average height of the population, why is she doing a hypothesis test that the height isn't 65 inches?

Since her hypothesis test is designed to help answer the question of whether or not the mean population height is 65 inches, why should we expect it to tell us anything about the mean population height other than whether or not it being 65 is consistent with the data observed?

Post reply on HN