Live data from Hacker News

It’s not just p=0.048 vs. p=0.052

statmodeling.stat.columbia.edu

1–10 of 92 posts

Re: It’s not just p=0.048 vs. p=0.052

#2
It looks like the blog author completely missed the point of the statistical significance discussion going on. Most first-tier journals in the social sciences have an acceptance rate of about 5%. At the margins, the differences between acceptance and rejection could be having one more statistical significance result in the table than the paper that was submitted right before or after yours.

The problem with a 0.048 and a 0.052 is not a mathematical one but an interpretation one. Reviewers are condition to be very skeptical of non-significant results and use “under power-ness” as a grounds for rejection. As a result, we get publication bias and p-hacking.

Re: It’s not just p=0.048 vs. p=0.052

#3
post #2

It looks like the blog author completely missed the point of the statistical significance discussion going on. Most first-tier journals in the social sciences have an acceptance rate of about 5%. At the margins, the differences between acceptance and rejection could be having one more statistical significance result in the table than the paper that was submitted right before or after yours. The problem with a 0.048 a…

[deleted]

Re: It’s not just p=0.048 vs. p=0.052

#4
post #2

It looks like the blog author completely missed the point of the statistical significance discussion going on. Most first-tier journals in the social sciences have an acceptance rate of about 5%. At the margins, the differences between acceptance and rejection could be having one more statistical significance result in the table than the paper that was submitted right before or after yours. The problem with a 0.048 a…

You point out p-values can be troubling because there's all sorts of bad incentives that lead to p-hacking and publication bias. The author points out that even if those bad incentives didn't exist, p-values aren't all that useful to begin with. That's not "missing the point", it's just pointing out a different aspect of the situation.

Re: It’s not just p=0.048 vs. p=0.052

#5
> To say it again: it is completely consistent with the null hypothesis to see p-values of 0.2 and 0.005 from two replications of the same damn experiment.

I don't really follow this. Could someone clarify what is meant here? At what point would this author say something is not consistent with the null hypothesis?

Re: It’s not just p=0.048 vs. p=0.052

#6
post #2

It looks like the blog author completely missed the point of the statistical significance discussion going on. Most first-tier journals in the social sciences have an acceptance rate of about 5%. At the margins, the differences between acceptance and rejection could be having one more statistical significance result in the table than the paper that was submitted right before or after yours. The problem with a 0.048 a…

There is whole group of people in social sciences who are pushing for abandoning the null hypothesis testing methods. For the reason you mentioned, I cannot believe changing the test (or threshold of the test) would solve this issue. You ask people to find something significant or fit a model to a data — and tell them that's what matters — and they'll do it, either intentionally or unintentionally.

Machine Learning will (have) the same issue if the only thing that matters is hitting a certain level of accuracy given your model and data. This has been observed in Kaggle competitions over and over, you ask a group of people to find the best fit, and they'll, by learning your train, validation and test datasets.

As mentioned, problem is not p-value, or null hypothesis testing, the problem is journals who promoted the wrong incentive, and educators who were not aware of the consequences and propagated the wrong incentive (interpretation) to students.

Re: It’s not just p=0.048 vs. p=0.052

#7
post #2

It looks like the blog author completely missed the point of the statistical significance discussion going on. Most first-tier journals in the social sciences have an acceptance rate of about 5%. At the margins, the differences between acceptance and rejection could be having one more statistical significance result in the table than the paper that was submitted right before or after yours. The problem with a 0.048 a…

I think you should reread the article because it's exactly what the blog author says. Blog author who btw is Andrew Gelman, not just some random guy on Medium, his blog is well well worth reading. Fighting bad stats in science is kind of his hobby/life mission.

Re: It’s not just p=0.048 vs. p=0.052

#8
post #5

> To say it again: it is completely consistent with the null hypothesis to see p-values of 0.2 and 0.005 from two replications of the same damn experiment. I don't really follow this. Could someone clarify what is meant here? At what point would this author say something is not consistent with the null hypothesis?

He is coming at that conclusion from a Bayesian point of view to statistics. He is seeing the p-value as a random variable that can take values from 0 to 1 and follows some distribution. Under these hypotheses, observing a p-value of 0.20 and 0.005 is completely reasonable even if unlikely. Those are just two draws from a random variable.

Edit. Under Bayesian statistics testing the null hypothesis is a moot point as it becomes possible to directly model the distribution of the possible effects. Thinking of it as being able to look at a picture of something (the p-value) vs looking at a movie of it (the distribution of the effects).

Re: It’s not just p=0.048 vs. p=0.052

#9
post #8
post #5

> To say it again: it is completely consistent with the null hypothesis to see p-values of 0.2 and 0.005 from two replications of the same damn experiment. I don't really follow this. Could someone clarify what is meant here? At what point would this author say something is not consistent with the null hypothesis?

He is coming at that conclusion from a Bayesian point of view to statistics. He is seeing the p-value as a random variable that can take values from 0 to 1 and follows some distribution. Under these hypotheses, observing a p-value of 0.20 and 0.005 is completely reasonable even if unlikely. Those are just two draws from a random variable. Edit. Under Bayesian statistics testing the null hypothesis is a moot point as…

What he says is (I gather) worse: those events are only separated be 1.1std deviations, which is little.

Re: It’s not just p=0.048 vs. p=0.052

#10
post #5

> To say it again: it is completely consistent with the null hypothesis to see p-values of 0.2 and 0.005 from two replications of the same damn experiment. I don't really follow this. Could someone clarify what is meant here? At what point would this author say something is not consistent with the null hypothesis?

.
Post reply on HN