Live data from Hacker News

Statisticians want to abandon science’s standard measure of ‘significance’

sciencenews.org

61–70 of 142 posts

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#61

If you're looking for a replacement you don't understand the problem. The problem isn't that P=.05 is an arbitrary measure of significance. The problem is that only publishing significant results is a bias against the null hypothesis . Let's say you're doing a study of flipping coins. The null hypothesis is that the coin is evenly weighted. If the null hypothesis is true, when you flip a coin once, it will come up he…

This is actually a much better explanation of the problem than the main article. The main article kept saying that p values were not meant to be definitive, but without explaining what was wrong with them. At least, not very clearly. This comment is a much clearer explanation -- imho -- as to what goes wrong in the industry. I also think it is right to focus the attention away from a particular statistical measure.

It seems clear to me that the author just didn't understand the problem with P values. It's part of a larger problem of science journalism being done by journalists without scientific backgrounds. It's not their fault--even if you have natural ability and interest in both communication and discovery, it's hard to get an education in both.

I have the opposite problem: my abilities lie more in the statistics/science than the communication--I suspect the only reason that this explanation is being received so well is that I happened across an effective example by pure luck. The probability of me communicating effectively is probably P < .5. ;)

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#62
post #56

Earlier quoted context omitted.

> Changing the nomenclature to ‘statistically discriminable’ would go a long way to improving popular understanding. I think you overestimate the intelligence of people prone to misunderstanding. I don't think they're likely to understand big words like "statistically" or "discriminable" when they already don't understand big words like "statistical" or "significance".

when they already don't understand big words like "statistical" or "significance". They do understand what "significant" means, they use the world every day. The problem is that it means something different than the intuitive every day meaning when used in context of p-values. Using a word like "discriminable" might help clear things up since it's a word that doesn't have have so much meaning packed into it already.…

> The problem is that it means something different than the intuitive every day meaning

Yes that's my point. Using words that are hard to understand or only understandable in the context of P-values won't solve the problem; the problem being that the general public won't understand the subtle differences in meaning

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#63
post #28

Earlier quoted context omitted.

It's worth to point that "3% probability we're just seeing a pattern by accident" is only right when you understand it as "in the world here our hypothesis is wrong the same experiment would give such pattern in 3% cases", not as "given such result probability that we are wrong is 3%".

It seems like an education fail to me. Most people don't know the very basics of stats and we live in a world that's highly probabilistic. It seems like something that should be taught alongside math from elementary school, not something you can maybe get an elective in in high school or college.

This guy is one of my favorite TED speakers, and he makes a strong argument for exactly this: https://www.ted.com/talks/arthur_benjamin_s_formula_for_chan....

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#64

The problem isn't p-values, the problem is a binary distinction between p=0.049 and p=0.051. The problem would go away if everyone understood p-values, or we replaced use of the term "statistically significant" with "3% probability we're just seeing a pattern by accident". Renaming the term to something that sounds just as binary isn't any different.

This is a a really common misconception about p values (that they can be interpreted as p(H0|x), or "probability of the null hypothesis given the data") when a p-value is in fact p(x|H0), or "probability of observing data at least this extreme given that the null hypothesis is true

Do you have a better of wording for inclusion in paper abstracts / article summaries than what I said? I'd love to hear one, but p(x|H0) is just as bad as p = 0.05.

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#65
I am pretty surprised to see no discussion of statistical power in the article and very little mention in the comments here. To me, having more statistical power solves many of the issues mentioned in the article. Many of the rest can be handled with use of Bayesian priors, context-specific p-values thresholds (.01, .1, etc.), and replication.

There's decent working guidelines for statistical power, but a lot of the issues (especially around replication) I see are mainly due to sample sizes being too small to adequately detect effects, especially if those effects are small.

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#66
post #43

The problem isn't p-values, the problem is a binary distinction between p=0.049 and p=0.051. The problem would go away if everyone understood p-values, or we replaced use of the term "statistically significant" with "3% probability we're just seeing a pattern by accident". Renaming the term to something that sounds just as binary isn't any different.

We use hard cutoffs for a bunch of things, they aren't perfect but they are fine. The problem is that we are imbuing the words "statistical significance" with a whole bunch of math. This would be fine, except for the inconvenient fact that people also want to use the word "significance" as it is defined in English. It is not only possible, but likely that people will be producing results that are insignificant but st…

This is a really good point.

In fact, failures to reproduce "landmark" experiments have low significance (statistics word) because they find a result that is highly probable if the null hypothesis is true. But failures to reproduce have high significance (colloquial English word) because they indicate that widely-held beliefs are wrong. So in some of the most important cases, the statistics and colloquial English meanings of the words are exactly the opposite.

I don't know if there's a good solution to this, though--it's hard to change organically-emergent trends in language even if we did come up with a better term.

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#67
I think the best term would have been statistically surprising, because it strongly hint at the fact that the result would be surprising under the null hypothesis, witch really is all that "statistically significant" really means. Sometimes surprising results happen, but all other things being equal they might hint at the null hypothesis being false. I could also live with "statistically interesting". "Detectable", suggested in another comment, seems to have some of the same issues as significant, it is too strong and seems to imply that now we know something is really there.

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#68
Although well written, this article misses the main point why statistical significance leads us astray and needs to be deemphasized. It is because "reproducibility" needs to be the new gold standard that displaces "significance testing". There are too many highly significant results that are totally unreproducible.

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#69

The problem isn't p-values, the problem is a binary distinction between p=0.049 and p=0.051. The problem would go away if everyone understood p-values, or we replaced use of the term "statistically significant" with "3% probability we're just seeing a pattern by accident". Renaming the term to something that sounds just as binary isn't any different.

At the end of the day, people need to make a decision about something, and I suspect this is what is forcing the importance of this binary distinction.

My decision to select, implement, go, or no go cannot itself be a probability distribution.

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#70
post #69

The problem isn't p-values, the problem is a binary distinction between p=0.049 and p=0.051. The problem would go away if everyone understood p-values, or we replaced use of the term "statistically significant" with "3% probability we're just seeing a pattern by accident". Renaming the term to something that sounds just as binary isn't any different.

At the end of the day, people need to make a decision about something, and I suspect this is what is forcing the importance of this binary distinction. My decision to select, implement, go, or no go cannot itself be a probability distribution.

Decision to select shouldn't be the subject of academic or scientific papers in the first place. There's far more context required in decision making than just how strong a correlation is.
Post reply on HN