Earlier quoted context omitted.
Thanks for the very interesting thoughts. Could you explain a bit how your rule of thumb works and why it's better than p-values? Why is the difference vs. the square root of the max available sample size a meaningful measure?
The idea is that you will decide when you've either expended as much energy as you are willing to, or when you're convinced that you won't make a different decision. There is a simple symmetry argument that shows that the odds of a random walk reaching sqrt(N) in one direction and then getting to the opposite direction by N observations is exactly the same as the odds of a random walk reaching 2 sqrt(N) by the time y…
Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values
101–110 of 130 posts
Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values
#102Earlier quoted context omitted.
But if they are having children until they have one of each, the probability of the different observations does depend on the prior events!
The absolute probability of the observation is irrelevant. Only the RELATIVE probabilities of said observation under the different possible theories which are part of the prior. If the set of prior theories does not include anything that depends on birth order, then birth order and the experimental design are irrelevant to the posterior conclusions.
Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values
#103For context, I have a degree in statistics, and I did research with Andrew Gelman (one of the statisticians quoted in the article). Glad to see this is gaining traction! I've been saying this for years: the world would actually be in a better place if we just abandoned p-values altogether. Hypothesis testing is taught in introductory statistics courses because the calculations involved are deceptively easy, whereas m…
If you're doing a measurement, why have a null hypothesis? Alice should sample the population at random, take the height measurements, calculate the average, plot the distribution, calculate the variance. If the distribution is not sufficiently smooth then continue to take measurements until it's smooth, or unchanging. Then Alice is done discovering all there is to know about the height distribution of the population…
Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values
#104Earlier quoted context omitted.
The absolute probability of the observation is irrelevant. Only the RELATIVE probabilities of said observation under the different possible theories which are part of the prior. If the set of prior theories does not include anything that depends on birth order, then birth order and the experimental design are irrelevant to the posterior conclusions.
Right, but this is exactly my point: the correct answer depends on the set of prior theories , which is exactly what the frequentist scenarios you consider are varying.
We have 3 types of facts.
1. What is the set of prior theories? Probability theory says this should matter. Bayesians are explicit about its involvement. Frequentists ignore it.
2. Observed data. Probability theory says that this should matter. Everyone takes this into account.
3. Experimental design for what would have happened had something different than the observed actually happened. This matters a lot to frequentist approaches and does not matter at all to Bayes' theorem. Bayesian approaches generally do not care about it at all.
The difference between scenario 1 and scenario 2 is a fact of type 3, the conditions under which Bill and Lorena would have stopped having children. Unless you believe that Bill and Lorena's desire for one gender has a material impact on the probability of boys vs girls, this fact is irrelevant to any calculation of posterior probabilities. And is irrelevant in classical Bayesian approaches.
Yet, despite being irrelevant, it matters a lot for frequentist approaches.
Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values
#105You can convert p-values to likelihood ratios, and they are quite similar. But its not perfect. A p value of 0.05 becomes 100:5, or 20:1. Which means it increases the odds of a hypothesis by 20. So a probability of 1% updates to 17%, which is still quite small.
But that assumes that the hypothesis has a 100% chance of producing the same or greater result, which is unlikely. Instead it might only be 50% which is half as much evidence.
In the extreme case, it could be only 5% likely to produce the result, which means the likelihood ratio is 5:5 and is literally no evidence, but still has a p value of 0.05.
Anyway likelihood ratios accumulate exponentially, since they multiply together. As long as there is no publication bias, you can take a few weak studies and produce a single very strong likelihood update.
Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values
#106Earlier quoted context omitted.
Right, but I guess what I'm getting at is frequently we see people doing even worse: making policy decisions based on data that doesn't even hit a minimum threshold for acceptability. I totally agree that people shouldn't make policy decisions solely because the data supports it, but frequently you see people make decisions based on data indicating something without actually being statistically significant. That feel…
One issue is that if you have a large effect that's consistently and easily reproduced, you don't actually need very accurate measurements or a statistical analysis at all. So any minimum standard would need to take into account. Another issue is that science is expensive and we need to make decisions all the time whether there is any science backing them or not. So what do you do if there's no science that meets the…
Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values
#107Earlier quoted context omitted.
Right, but I guess what I'm getting at is frequently we see people doing even worse: making policy decisions based on data that doesn't even hit a minimum threshold for acceptability. I totally agree that people shouldn't make policy decisions solely because the data supports it, but frequently you see people make decisions based on data indicating something without actually being statistically significant. That feel…
One issue is that if you have a large effect that's consistently and easily reproduced, you don't actually need very accurate measurements or a statistical analysis at all. So any minimum standard would need to take into account. Another issue is that science is expensive and we need to make decisions all the time whether there is any science backing them or not. So what do you do if there's no science that meets the…
Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values
#108Earlier quoted context omitted.
Right, but this is exactly my point: the correct answer depends on the set of prior theories , which is exactly what the frequentist scenarios you consider are varying.
If you are arguing THAT point, then you shouldn't have disagreed with me anywhere! We have 3 types of facts. 1. What is the set of prior theories? Probability theory says this should matter. Bayesians are explicit about its involvement. Frequentists ignore it. 2. Observed data. Probability theory says that this should matter. Everyone takes this into account. 3. Experimental design for what would have happened had so…
Anyway, not sure if we're on the same page, but thanks for the discussion.
Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values
#109If they held the same meeting 20 times, would they reach the same conclusion in 19 of those meetings? On a more serious note, I think that the use of the word "significant" to mean "the effect is reasonably likely to exist by some standard" should be abolished. Webster's 1913 dictionary says: > Deserving to be considered; important; momentous; as, a significant event. Statisticians don't use "significant" to mean imp…
"Eating fat had no discernible effect on weight gain" sounds like getting evidence against such an effect, but it's compatible with getting evidence in favor, that's just not as strong as some threshold. That evidence could be useful in a meta-analysis, or for a decision when waiting for more information isn't practical or economical, or if the potential gain from trying the nonsignificant treatment is high and the potential loss low. (I've seen "no significant X" abused this way. Nobody should try X -- it's unscientific!)
Re: Statisticians Find They Can Agree: It’s Time to Stop Misusing P-Values
#110Earlier quoted context omitted.
If you are arguing THAT point, then you shouldn't have disagreed with me anywhere! We have 3 types of facts. 1. What is the set of prior theories? Probability theory says this should matter. Bayesians are explicit about its involvement. Frequentists ignore it. 2. Observed data. Probability theory says that this should matter. Everyone takes this into account. 3. Experimental design for what would have happened had so…
1 and 3 are really the same thing. If the complaint is that the frequentist test can't tell you anything if the assumed distribution wasn't the right one (which is what's happening if you would have done something different), consider the bayesian case. There one might argue you at least still have the probability of each hypothesis given the data. But that forgets that it is only the probability of each hypothesis g…
In particular 1 consists of exact statements about the likelihood of 7 births in a row being mmmmmmf.
By contrast 3 consists of statements about what Bill and Lorena's childbearing plans would have been if something different had happened.
Those are very different types of statement. There is no connection between statements of type 3 and statements of type 1 unless Bill and Lorena's state of mind makes a significant difference to the odds of the next child being a boy.