Live data from Hacker News

Statisticians want to abandon science’s standard measure of ‘significance’

sciencenews.org

71–80 of 142 posts

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#71
post #43

The problem isn't p-values, the problem is a binary distinction between p=0.049 and p=0.051. The problem would go away if everyone understood p-values, or we replaced use of the term "statistically significant" with "3% probability we're just seeing a pattern by accident". Renaming the term to something that sounds just as binary isn't any different.

We use hard cutoffs for a bunch of things, they aren't perfect but they are fine. The problem is that we are imbuing the words "statistical significance" with a whole bunch of math. This would be fine, except for the inconvenient fact that people also want to use the word "significance" as it is defined in English. It is not only possible, but likely that people will be producing results that are insignificant but st…

This is an excellent point. In my work, the interpretation error I see the most among people around be is to confuse significance with the magnitude of the predicted effect. I.e. if you use OLS to create a model, it is possible to have a statistically significant beta, with essentially no magnitude.

If you look at research on carcinogens, it is frequent to see something like "it is statistically significant that eating chemical X increases your risk of cancer" but when you read on, the magnitude of the effect is "increases risk by .0023%, for a cancer that already has a .6% incidence in the population".

My point is that people who are not statisticians (or people who use statistics), tend to be concerned with the magnitude of an effect, and frequently misconstrue the statistical significance for that magnitude.

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#72
post #69

Earlier quoted context omitted.

At the end of the day, people need to make a decision about something, and I suspect this is what is forcing the importance of this binary distinction. My decision to select, implement, go, or no go cannot itself be a probability distribution.

Decision to select shouldn't be the subject of academic or scientific papers in the first place. There's far more context required in decision making than just how strong a correlation is.

I agree, but you might see how this error or preference creeps in psychologically

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#73

Earlier quoted context omitted.

This is a a really common misconception about p values (that they can be interpreted as p(H0|x), or "probability of the null hypothesis given the data") when a p-value is in fact p(x|H0), or "probability of observing data at least this extreme given that the null hypothesis is true

Do you have a better of wording for inclusion in paper abstracts / article summaries than what I said? I'd love to hear one, but p(x|H0) is just as bad as p = 0.05.

The problem isn't that the wording is confusing, it's that p(x|H0) isn't a very useful thing to compute.

Everybody wants to compute p(H|x), the probility of a scientific hypothesis given the data. People want to do this so badly that they can't help interpreting the p-value that way.

You can actually compute p(H|x) if you use Bayesian stats.

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#76

Earlier quoted context omitted.

Do you have a better of wording for inclusion in paper abstracts / article summaries than what I said? I'd love to hear one, but p(x|H0) is just as bad as p = 0.05.

The problem isn't that the wording is confusing, it's that p(x|H0) isn't a very useful thing to compute. Everybody wants to compute p(H|x), the probility of a scientific hypothesis given the data. People want to do this so badly that they can't help interpreting the p-value that way. You can actually compute p(H|x) if you use Bayesian stats.

With some caveats: the value you get for P(H|x) might be very sensitive to the priors you choose, and most people are not thoughtful enough about this.

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#77

The problem isn't p-values, the problem is a binary distinction between p=0.049 and p=0.051. The problem would go away if everyone understood p-values, or we replaced use of the term "statistically significant" with "3% probability we're just seeing a pattern by accident". Renaming the term to something that sounds just as binary isn't any different.

I disbelieve.

The problem is that people want an answer to the question, "Here is a pile of data, what should I believe?" But mathematically, the proper role of data is to modify existing beliefs, and not to dictate beliefs.

Statisticians can spend forever explaining this. But instead have gone with the cop-out of asking a question that is confusingly similar to the one that people want to ask. It is popular exactly because it is so easily misunderstood. You can give any number of lectures on what it actually means - I guarantee that it will be misunderstood.

Worse yet, p-values are sensitive to particulars of experimental design that logically should never matter to inferences. My aunt and uncle provide a classic example. They wanted a son and a daughter. They had 6 sons then a daughter. Are they biased towards one gender? The p-value for this question is .03125 (we'd see this much evidence against with all boys, all girls, last girl, last boy, so 4/2^7). But if I instead said that they planned to have 7 children, and had 6 sons then a daughter, the p-value changes to 0.125 because the one child out could have been anywhere in succession. But in Bayes' formula, their plans if something else had happened cannot ever affect how we adjust our inferences.

(This example is not made up. When Lorna found that they had a daughter, she told Bill to get a vasectomy from the delivery table. By another crazy coincidence, their 7 children were also born on all the days of the week.)

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#78
post #65

I am pretty surprised to see no discussion of statistical power in the article and very little mention in the comments here. To me, having more statistical power solves many of the issues mentioned in the article. Many of the rest can be handled with use of Bayesian priors, context-specific p-values thresholds (.01, .1, etc.), and replication. There's decent working guidelines for statistical power, but a lot of the…

Can you recommend any texts, videos, or websites to learn more about statistical power? The term is quite overloaded in google.

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#79
post #77

The problem isn't p-values, the problem is a binary distinction between p=0.049 and p=0.051. The problem would go away if everyone understood p-values, or we replaced use of the term "statistically significant" with "3% probability we're just seeing a pattern by accident". Renaming the term to something that sounds just as binary isn't any different.

I disbelieve. The problem is that people want an answer to the question, "Here is a pile of data, what should I believe?" But mathematically, the proper role of data is to modify existing beliefs, and not to dictate beliefs. Statisticians can spend forever explaining this. But instead have gone with the cop-out of asking a question that is confusingly similar to the one that people want to ask. It is popular exactly…

Well, sort-of.

Its not just about p-values or the likelyhood of making a particular type of error by agreeing with a thesis.

There is a fundamental misunderstanding that most people, and even many scientists make with regards to the philosophical interpretation of 'truth' using the scientific method.

By definition, scientific 'truth' is always fallible, and the vast majority of people have a very hard time dealing with this concept. A better explanation or additional data, or better data is always possible and needs to be possible for science to not be dogma. This is fundamentally different than how people interpret truth and understand what the word means. All scientific truths are replaceable.

Re: Statisticians want to abandon science’s standard measure of ‘significance’

#80

There's no way to guard against all false positives. And while changing the P value cutoff would reduce them, it would also increase false negatives. The answer doesn't lie in hard-line stances for or against P values or with an alternative that will have its own set of problems. It lies with greater education of those who run experiments & those who consumer the literature about proper interpretation and other metho…

Other small things could help, like publishing failures, incentivizing duplication of experiments, increasing access to results.

I think duplication and access to raw data would be two huge steps forward. It's actually a bit mystifying that researchers are able to present their own final conclusions & interpretations without making the actual data available.
Post reply on HN