Live data from Hacker News

Statisticians want to abandon science’s standard measure of ‘significance’

sciencenews.org

111–120 of 142 posts

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#111

If you're looking for a replacement you don't understand the problem. The problem isn't that P=.05 is an arbitrary measure of significance. The problem is that only publishing significant results is a bias against the null hypothesis . Let's say you're doing a study of flipping coins. The null hypothesis is that the coin is evenly weighted. If the null hypothesis is true, when you flip a coin once, it will come up he…

Is there a compendium somewhere of null hypotheses as well as observed p values (or whatever statistic) for experiments with both "significant" and "insignificant" results? Depending on the phenomenon and the hypothesis, an event could have a different probability of occurring and require a different threshold.

It seems like we are wasting a lot of effort when any experiment is unpublished, when we could at the least be associating a data point with a particular hypothesis.

This would require some standardization of hypotheses, so that a researcher could select a hypothesis from a list, conduct an experiment to test it, and report those findings to some aggregator. Others would also test the hypothesis and report their findings. Eventually you have some distribution of findings that allow you to examine the experimental methods of outliers as well as modal experiments. In this way every scientific result is the product of some meta-analysis, rather than allowing a single custom experiment to produce a result.

The current approach also ties the experimental method to the result, requiring potentially fallible oversight of the experimental design. The hypothesis-testing aggregator reduces this reliance, which is perhaps more honest about our ability to design good experiments, would allow for a greater diversity of approaches, and could also be more efficient when evaluating results.

The most interesting aspect of this is if you could somehow create relationships between standardized hypotheses found on the hypothesis menu such that you could infer other hypotheses in a more systematic way.

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#112
post #67

I think the best term would have been statistically surprising, because it strongly hint at the fact that the result would be surprising under the null hypothesis, witch really is all that "statistically significant" really means. Sometimes surprising results happen, but all other things being equal they might hint at the null hypothesis being false. I could also live with "statistically interesting". "Detectable", s…

By the reasoning behind significance, "surprising" would be a great drop-in. However, in most studies, it would be more surprising if the null hypothesis were true. Statistically significant results are pretty much a given.

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#113

Earlier quoted context omitted.

The problem isn't that the wording is confusing, it's that p(x|H0) isn't a very useful thing to compute. Everybody wants to compute p(H|x), the probility of a scientific hypothesis given the data. People want to do this so badly that they can't help interpreting the p-value that way. You can actually compute p(H|x) if you use Bayesian stats.

> You can actually compute p(H|x) if you use Bayesian stats. You can't "compute" it for any useful meaning of the word "compute". You can estimate it intuitively, or you can try to look at how many pre-registered unpublished studies or null results have been published. Otherwise there's no way to even consider getting a grasp of what P(H) would be.

> Otherwise there's no way to even consider getting a grasp of what P(H) would be.

That's called picking a prior. Usually, pick the prior that maximizes the entropy on the given interval (zero prior knowledge).

For those reading and confused. Computing: p(H|x) = p(x|H)p(H)/p(x)

requires p(H) which the parent is suggesting is impossible to grasp.

If we want to know the probability that a coin is biased, we can assign probabilities to each hypothesis. For some people, p(bias=1/2)=1 (Dirac distribution), others might argue that p(bias=x)=1 for x in [0,1] (uniform distribution). Others might argue its some Beta function centered around 1/2. I believe what the parent is suggesting is that choosing which original belief we have in the system is a matter of philosophy, not computation.

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#114
I'm not a statistician, just a biologist. But I have seen some pretty low P-values indicating two distribution have a differing mean while the plot right on top of each other. This is when there are a lot of measurements (in my case there were over 100.000). I prefer to make ROC curves or simple just to look at the plotted distributions. ROC curves give a nice idea about the mixed-ness of positive and negative measurements.

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#115

The problem isn't p-values, the problem is a binary distinction between p=0.049 and p=0.051. The problem would go away if everyone understood p-values, or we replaced use of the term "statistically significant" with "3% probability we're just seeing a pattern by accident". Renaming the term to something that sounds just as binary isn't any different.

I’m a Bayesian flavoured, and the “problem” IMO is that model driven statistics are really hard and beyond most scientists (people generally tbh). If we required scientists to be good statisticians there’d be far fewer scientists.

Or we'd have just as many, but better at statistics.

P hacking isn't always done on purpose. It's misunderstanding. That's probably why it still passes peer review. There's also tons of incentives that encourage this behavior. These are the problems. An arbitrary mark to meet makes people weak at statistics not understand their data as well. I'd rather more accurate data than more scientists. More workers doing poor work isn't useful.

Also, you can Bayes hack. Bayes helps, but it doesn't address the underlying issues.

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#116

Most comments here point to cherry picking and "p hacking" as being the primary problems with p values. Certainly those are major issues, but I think they miss the real point of the article, which is that null hypothesis testing is fundamentally broken, or at the very least doesn't do what most people think. A simple example of this can be shown with the following pair of tests: Testing for a fair coin: - Null hypoth…

The problem in the second case is not null-hypothesis testing, it’s sample size.

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#117
p=0.05 is an absurdly low bar even with an extremely well-defined hypothesis. Once you open things to p-value hacking (i.e. trying a bunch of hypotheses until one seems significant), p=0.05 is almost guaranteed for one of your hypotheses and is hence meaningless.

The physics community is pretty good about dealing with both of these shortcomings. You often need a lower p-value of around p=0.003 (3 sigma) or lower to say something was significant, and a few more sigmas to claim an actual discovery. And in addition to this, you're expected to include a trials factor to correct for multiple hypotheses. It's not perfect, but it makes claims of statistical significance more meaningful.

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#118

Earlier quoted context omitted.

This is a a really common misconception about p values (that they can be interpreted as p(H0|x), or "probability of the null hypothesis given the data") when a p-value is in fact p(x|H0), or "probability of observing data at least this extreme given that the null hypothesis is true

This is a more mathematically precise restatement of kerkeslager's explanation below and the obligatory xkcd [1]. The problem boils down to the fact that null results are not usually published, and that intuitive skepticism is not quantifiable. Essentially, we want to get p(H0|x), but we need Bayes Law to get this from p(x|H0). But we need some notion of what priors to use. This is of course impossible to actually ge…

I'd argue that p(H0|x) is also (in most cases) pretty uninteresting from a scientific perspective, and so the whole "publish all the p-values" solution I think is only fighting half the battle. Gelman does a much better job of arguing this than I ever could [1], but this idea that rejecting the null is the "desired" outcome is the real problem. To put it another way, a low p-value is another way of saying "my model of the data generating process is bad", which is in most cases a pretty unsatisfactory outcome. As to the idea that priors are "impossible" to get, to quote Gelman [2] again, why "strain at the gnat that is the prior and swallow the ungainly camel that is the iid likelihood?

[1] https://statmodeling.stat.columbia.edu/2015/03/02/what-hypot... [2] https://statmodeling.stat.columbia.edu/2015/07/03/why-should...

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#119

If you're looking for a replacement you don't understand the problem. The problem isn't that P=.05 is an arbitrary measure of significance. The problem is that only publishing significant results is a bias against the null hypothesis . Let's say you're doing a study of flipping coins. The null hypothesis is that the coin is evenly weighted. If the null hypothesis is true, when you flip a coin once, it will come up he…

> we can conclude that the coin is weighted in some way with P=0.03125

No we can't. There are 2 (related) reasons for this. The first is that to say "after experiment X, we can conclude thing Y is happening with probability ..." you need to know something about how often thing Y happens on its own. This is also known as a prior. The other reason is that in this specific case, the prior for Y (that the coin is weighted) can be concisely summarized as P(Y)=0, because it's physically impossible to construct a coin that's biased to one side or the other (independent of what side it starts on). Some people think the coin flipping examples are bad because biased coins are impossible, but they're actually good for exactly this reason. We have a really good prior (from math/physics/geometry) that biased coins can't exist, so you're going to have to not only flip a lot of them to convince me otherwise, but also give me an explanation of how that's possible (i.e justify your prior)

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#120

This is a very old argument. I got my bachelor's degree in psychology at Harvard in 1993, and was told repeatedly that p-tests are abused, overused, and not terribly useful. To my mind, the most hackable flaw is that the number of subjects in the study is a term in the denominator of the p-value calculation. Any study with a sufficiently large sample will find "significance" with p We were taught that "effect size" m…

> To my mind, the most hackable flaw is that the number of subjects in the study is a term in the denominator of the p-value calculation. Any study with a sufficiently large sample will find "significance" with pJust to clarify. Any study with a sufficiently large sample will find a real effect even if the effect isn't clinically significant.
Post reply on HN