Live data from Hacker News

Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

fivethirtyeight.com

101–110 of 130 posts

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#101
post #58

Earlier quoted context omitted.

Thanks for the very interesting thoughts. Could you explain a bit how your rule of thumb works and why it's better than p-values? Why is the difference vs. the square root of the max available sample size a meaningful measure?

The idea is that you will decide when you've either expended as much energy as you are willing to, or when you're convinced that you won't make a different decision. There is a simple symmetry argument that shows that the odds of a random walk reaching sqrt(N) in one direction and then getting to the opposite direction by N observations is exactly the same as the odds of a random walk reaching 2 sqrt(N) by the time y…

Thanks! That's very helpful.

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#102
post #98

Earlier quoted context omitted.

But if they are having children until they have one of each, the probability of the different observations does depend on the prior events!

The absolute probability of the observation is irrelevant. Only the RELATIVE probabilities of said observation under the different possible theories which are part of the prior. If the set of prior theories does not include anything that depends on birth order, then birth order and the experimental design are irrelevant to the posterior conclusions.

Right, but this is exactly my point: the correct answer depends on the set of prior theories, which is exactly what the frequentist scenarios you consider are varying.

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#103

For context, I have a degree in statistics, and I did research with Andrew Gelman (one of the statisticians quoted in the article). Glad to see this is gaining traction! I've been saying this for years: the world would actually be in a better place if we just abandoned p-values altogether. Hypothesis testing is taught in introductory statistics courses because the calculations involved are deceptively easy, whereas m…

If you're doing a measurement, why have a null hypothesis? Alice should sample the population at random, take the height measurements, calculate the average, plot the distribution, calculate the variance. If the distribution is not sufficiently smooth then continue to take measurements until it's smooth, or unchanging. Then Alice is done discovering all there is to know about the height distribution of the population…

Creating an 100(1-a)% interval estimate of the population mean based on the sample's mean and variance is equivalent to examining the set of all values which produce a p-value below a%. Your view of the way Alice and Bob should be making inference about the true mean is actually exactly what hypothesis testing is. It's just that there is often a "status quo" belief about the parameter which has important theoretical implications. And academic papers are usually about that belief specifically, so it makes sense that they talk about it as a "hypothesis test". And that's the logical headline, the statement about the specific value. But standard deviations and distributions are often given, making it trivial to create the interval as a reader.

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#104
post #98

Earlier quoted context omitted.

The absolute probability of the observation is irrelevant. Only the RELATIVE probabilities of said observation under the different possible theories which are part of the prior. If the set of prior theories does not include anything that depends on birth order, then birth order and the experimental design are irrelevant to the posterior conclusions.

Right, but this is exactly my point: the correct answer depends on the set of prior theories , which is exactly what the frequentist scenarios you consider are varying.

If you are arguing THAT point, then you shouldn't have disagreed with me anywhere!

We have 3 types of facts.

1. What is the set of prior theories? Probability theory says this should matter. Bayesians are explicit about its involvement. Frequentists ignore it.

2. Observed data. Probability theory says that this should matter. Everyone takes this into account.

3. Experimental design for what would have happened had something different than the observed actually happened. This matters a lot to frequentist approaches and does not matter at all to Bayes' theorem. Bayesian approaches generally do not care about it at all.

The difference between scenario 1 and scenario 2 is a fact of type 3, the conditions under which Bill and Lorena would have stopped having children. Unless you believe that Bill and Lorena's desire for one gender has a material impact on the probability of boys vs girls, this fact is irrelevant to any calculation of posterior probabilities. And is irrelevant in classical Bayesian approaches.

Yet, despite being irrelevant, it matters a lot for frequentist approaches.

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#105
P-values are so weird. Studies should instead report a likelihood ratio. A likelihood ratio is mathematically correct, and tells you exactly how much to update a hypothesis.

You can convert p-values to likelihood ratios, and they are quite similar. But its not perfect. A p value of 0.05 becomes 100:5, or 20:1. Which means it increases the odds of a hypothesis by 20. So a probability of 1% updates to 17%, which is still quite small.

But that assumes that the hypothesis has a 100% chance of producing the same or greater result, which is unlikely. Instead it might only be 50% which is half as much evidence.

In the extreme case, it could be only 5% likely to produce the result, which means the likelihood ratio is 5:5 and is literally no evidence, but still has a p value of 0.05.

Anyway likelihood ratios accumulate exponentially, since they multiply together. As long as there is no publication bias, you can take a few weak studies and produce a single very strong likelihood update.

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#106
post #49

Earlier quoted context omitted.

Right, but I guess what I'm getting at is frequently we see people doing even worse: making policy decisions based on data that doesn't even hit a minimum threshold for acceptability. I totally agree that people shouldn't make policy decisions solely because the data supports it, but frequently you see people make decisions based on data indicating something without actually being statistically significant. That feel…

One issue is that if you have a large effect that's consistently and easily reproduced, you don't actually need very accurate measurements or a statistical analysis at all. So any minimum standard would need to take into account. Another issue is that science is expensive and we need to make decisions all the time whether there is any science backing them or not. So what do you do if there's no science that meets the…

in these responses that is implied and a view inherently held by many: P-Value Science. != not true. A P-Value is just one way, and usually not a very good one. For most business decisions, median/mean (depending on data) absolute error in hold out is great.

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#107
post #49

Earlier quoted context omitted.

Right, but I guess what I'm getting at is frequently we see people doing even worse: making policy decisions based on data that doesn't even hit a minimum threshold for acceptability. I totally agree that people shouldn't make policy decisions solely because the data supports it, but frequently you see people make decisions based on data indicating something without actually being statistically significant. That feel…

One issue is that if you have a large effect that's consistently and easily reproduced, you don't actually need very accurate measurements or a statistical analysis at all. So any minimum standard would need to take into account. Another issue is that science is expensive and we need to make decisions all the time whether there is any science backing them or not. So what do you do if there's no science that meets the…

This is why the phrase 'An anecdote is not data' chaps my hide. If you have good observability and theory to back it up a single measurement can tell you everything. Where when you don't have that, take all the data points you want. All you have is garbage.

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#108
post #104

Earlier quoted context omitted.

Right, but this is exactly my point: the correct answer depends on the set of prior theories , which is exactly what the frequentist scenarios you consider are varying.

If you are arguing THAT point, then you shouldn't have disagreed with me anywhere! We have 3 types of facts. 1. What is the set of prior theories? Probability theory says this should matter. Bayesians are explicit about its involvement. Frequentists ignore it. 2. Observed data. Probability theory says that this should matter. Everyone takes this into account. 3. Experimental design for what would have happened had so…

1 and 3 are really the same thing. If the complaint is that the frequentist test can't tell you anything if the assumed distribution wasn't the right one (which is what's happening if you would have done something different), consider the bayesian case. There one might argue you at least still have the probability of each hypothesis given the data. But that forgets that it is only the probability of each hypothesis given the data, given that those were the only possible hypotheses. And if you admit the possibility of their being other possible hypotheses, then these probabilities don't really tell you anything either.

Anyway, not sure if we're on the same page, but thanks for the discussion.

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#109
post #21

If they held the same meeting 20 times, would they reach the same conclusion in 19 of those meetings? On a more serious note, I think that the use of the word "significant" to mean "the effect is reasonably likely to exist by some standard" should be abolished. Webster's 1913 dictionary says: > Deserving to be considered; important; momentous; as, a significant event. Statisticians don't use "significant" to mean imp…

While I agree that 'significant' has a misleading connotation, 'discernible' is also misleading. 'Statistically significant' just isn't any everyday concept, and trying to phrase it as one will encourage people to make mistakes. It's a complex concept: if we tried this experiment under a certain null hypothesis, then it'd be at least this improbable to see a result at least this extreme. The most I'd be willing to cut it down, after so much confusion in its actual use, is "subjunctively improbable", with the null hypothesis and the threshold left implicit. "Eating fat had no subjunctively improbable effect on weight gain." This sounds technical and fiddly, which I think is a feature: if you don't like it, don't base your reporting on a technical, fiddly concept.

"Eating fat had no discernible effect on weight gain" sounds like getting evidence against such an effect, but it's compatible with getting evidence in favor, that's just not as strong as some threshold. That evidence could be useful in a meta-analysis, or for a decision when waiting for more information isn't practical or economical, or if the potential gain from trying the nonsignificant treatment is high and the potential loss low. (I've seen "no significant X" abused this way. Nobody should try X -- it's unscientific!)

Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values

#110
post #104

Earlier quoted context omitted.

If you are arguing THAT point, then you shouldn't have disagreed with me anywhere! We have 3 types of facts. 1. What is the set of prior theories? Probability theory says this should matter. Bayesians are explicit about its involvement. Frequentists ignore it. 2. Observed data. Probability theory says that this should matter. Everyone takes this into account. 3. Experimental design for what would have happened had so…

1 and 3 are really the same thing. If the complaint is that the frequentist test can't tell you anything if the assumed distribution wasn't the right one (which is what's happening if you would have done something different), consider the bayesian case. There one might argue you at least still have the probability of each hypothesis given the data. But that forgets that it is only the probability of each hypothesis g…

We are clearly not on the same page. Because I think that 1 and 3 are rather different things, and you don't.

In particular 1 consists of exact statements about the likelihood of 7 births in a row being mmmmmmf.

By contrast 3 consists of statements about what Bill and Lorena's childbearing plans would have been if something different had happened.

Those are very different types of statement. There is no connection between statements of type 3 and statements of type 1 unless Bill and Lorena's state of mind makes a significant difference to the odds of the next child being a boy.

Post reply on HN