Live data from Hacker News

Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

fivethirtyeight.com

41–50 of 130 posts

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#41
post #6

I agree as well! Here is what probability theory teaches us. The proper role of data is to adjust our prior beliefs about probabilities to posterior beliefs through Bayes' theorem. The challenge is how to best communicate this result to people who may have had a wide range of prior beliefs. p-values capture a degree of surprise in the result. Naively, a surprising result should catch our attention and cause us to ret…

Simple Bayesian approaches take the opposite approach. You generally start with some relatively naive prior, and then treat the posterior as being the conclusion. Which is not very realistic if the real prior was something quite different.

I don't think this is a completely accurate portrayal of Bayesian stats. In Bayesian stats, there is no "real prior". Probability distributions are all subjective representations of belief. The prior is just what you believe prior to evidence, and the posterior is what you believe after you've taken evidence into account.

That said, moving away from p-values and towards something more robust is something the A/B testing industry needs. (Obviously I have my own opinion of what that something should be, and it's a bit different from what you are advocating.) There are far too many consultancies and agencies p-hacking their way to positive results ("hey unsophisticated client - guess what I made your conversion rate go up 25%!") and I'd love to see every one of them die.

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#42
post #20

For context, I have a degree in statistics, and I did research with Andrew Gelman (one of the statisticians quoted in the article). Glad to see this is gaining traction! I've been saying this for years: the world would actually be in a better place if we just abandoned p-values altogether. Hypothesis testing is taught in introductory statistics courses because the calculations involved are deceptively easy, whereas m…

For giggles and grins, my aunt and uncle are an actual example of one of the classic frequentist vs bayesian examples where frequentist statistics says that something utterly irrelevant should matter. Scenario 1 (true). Bill and Lorena had 7 children. 6 were boys, 1 was a girl. Are they biased towards having one gender? A 2-sided p-value says that there are 16 possibilities of this strength or more, each of which has…

Is this actually true for the Bayesian model? Throw in a parameter for whether or not Bill and Lorena were trying to have children until they had a boy and a girl. Now your answer will depend quite heavily on your prior!

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#43
post #39

p-value analysis has its big caveats like multiple comparisons, but Bayesian has its own, such as it's extremely hard to calculate priors. Both are challenging to use in difficult analysis and both can be abused.

> such as it's extremely hard to calculate priors. I think you mean it is difficult to formulate priors? Typically, calculating a prior only involves sampling from a distribution.

Though technically, obtaining that distribution, if you're doing anything above "Eh, it's probably normal..." could be described as "hard".

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#44

5% (1 in 20) is a pretty weak threshold to pass. Let's go 5 sigma (p < 3e-7) for discoveries and reserve 0.05 < p < 3e-7 for stuff we should take closer looks at.

I once obtained a p-value of zero (or more accurately, smaller than the numerical precision of a p-value in R) for a result that was, by design, meaningless.

It's a bad idea to use p-values as thresholds, regardless of where you put the threshold.

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#45

Earlier quoted context omitted.

Multiple comparisons? 1 million independent tests? Hello? >Your P value is 0.0002. This doesn't count as discovery, but clinically I know what my judgment is going to be. >Let's go 5 sigma (p ^Means you should start a new trial with n > 20 given the same placebo/drug split.

In the scientific world, it is economically infeasable (cost/time) for 1 million tests for a given hypothesis. Even n > 20 can be difficult for certain studies. Bootstrapping the results to simulate 1 million trials won't fix the aforementioned issue either.

> Even n > 20 can be difficult for certain studies.

For some expensive experiments even n=3 is difficult to reach. "Should we do another replicate and go from n=2 to n=3, or should we hire another post-doc?" is a relatively common (rhetorical) question.

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#46
post #20

For context, I have a degree in statistics, and I did research with Andrew Gelman (one of the statisticians quoted in the article). Glad to see this is gaining traction! I've been saying this for years: the world would actually be in a better place if we just abandoned p-values altogether. Hypothesis testing is taught in introductory statistics courses because the calculations involved are deceptively easy, whereas m…

For giggles and grins, my aunt and uncle are an actual example of one of the classic frequentist vs bayesian examples where frequentist statistics says that something utterly irrelevant should matter. Scenario 1 (true). Bill and Lorena had 7 children. 6 were boys, 1 was a girl. Are they biased towards having one gender? A 2-sided p-value says that there are 16 possibilities of this strength or more, each of which has…

Is this oddity just because scenario 1 fails to account for order?

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#47
post #30
post #6

I agree as well! Here is what probability theory teaches us. The proper role of data is to adjust our prior beliefs about probabilities to posterior beliefs through Bayes' theorem. The challenge is how to best communicate this result to people who may have had a wide range of prior beliefs. p-values capture a degree of surprise in the result. Naively, a surprising result should catch our attention and cause us to ret…

The way I see p-values is a mathematical way to share how certain you are that something is in the realm of reality (not truth). It is not a formal conclusion of the results, and is nothing more than a "spoiler" indicating what you might also conclude.

This point of view is not supported by the mathematics. If you think otherwise, then you do not understand the mathematics.

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#48
post #46
post #20

Earlier quoted context omitted.

For giggles and grins, my aunt and uncle are an actual example of one of the classic frequentist vs bayesian examples where frequentist statistics says that something utterly irrelevant should matter. Scenario 1 (true). Bill and Lorena had 7 children. 6 were boys, 1 was a girl. Are they biased towards having one gender? A 2-sided p-value says that there are 16 possibilities of this strength or more, each of which has…

Is this oddity just because scenario 1 fails to account for order?

Exactly. There are more ways to get this unusual outcome in scenario 1 than scenario 2.

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#49
post #37

Earlier quoted context omitted.

> "3. Scientific conclusions and business or policy decisions should not be based only on whether a p-value passes a specific threshold. There is no lower threshold at which the data becomes non-predictive?

More likely they mean the opposite: that a statistically significant P value (by whatever threshold you decide to use) should not be used, by itself, to drive policy decisions. Internally, the effect size still matters. Externally, there are numerous other factors that should drive decisionmaking.

Right, but I guess what I'm getting at is frequently we see people doing even worse: making policy decisions based on data that doesn't even hit a minimum threshold for acceptability. I totally agree that people shouldn't make policy decisions solely because the data supports it, but frequently you see people make decisions based on data indicating something without actually being statistically significant. That feels like a bigger problem to me than people actually getting good data but then using it too confidently, which I have rarely if ever seen.

TLDR: People tend to make policy decisions based on data even when the data is basically useless, I'm more worried about that then blindly following good data.

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#50

Earlier quoted context omitted.

If you're doing a measurement, why have a null hypothesis? Alice should sample the population at random, take the height measurements, calculate the average, plot the distribution, calculate the variance. If the distribution is not sufficiently smooth then continue to take measurements until it's smooth, or unchanging. Then Alice is done discovering all there is to know about the height distribution of the population…

> Alice should sample the population at random, take the height measurements, calculate the average, plot the distribution, calculate the variance A student-t test is basically a lossy encoding of that information you describe (the sample mean, the variance, and the sample size). > If the distribution is not sufficiently smooth then continue to take measurements until it's smooth, or unchanging. From a frequentist st…

[deleted]
Post reply on HN