Live data from Hacker News

Effect size is significantly more important than statistical significance

argmin.net

71–80 of 168 posts

Re: Effect size is significantly more important than statistical significance

#71

Earlier quoted context omitted.

The problem is that it’s really hard to get good data, ethically, in medical sciences. Something that improves outcomes by 5-10% can be really important, but trying to get a study big enough to prove it can be super expensive already.

Nobody likes being in the control group of the first working anti-aging serum...

> Nobody likes being in the control group of the first working anti-aging serum...

You only know whether it works when the study has been completed. You also only know whether the drug has (potentially) disastrous consequences when the study has been completed. Thus, I am not completely sure whether your claim holds.

Re: Effect size is significantly more important than statistical significance

#73
post #67

Earlier quoted context omitted.

And if misclassification is a concern (as the parent mentioned) you can put a prior on that rate too!

Which rate? The rate you failed to mix the balls? The rate you failed to count a ball? The rate you misclassified the ball? The rate you repeatedly counted the same ball? The rate you started with an incorrect count? The rate you did the math wrong? etc Here’s the experiment and here’s the data is concrete it may be bogus but it’s information. Updating probabilistic based on recursive estimates of probabilities is la…

> Which rate? The rate you failed to mix the balls? The rate you failed to count a ball? The rate you misclassified the ball? The rate you repeatedly counted the same ball? The rate you started with an incorrect count? The rate you did the math wrong? etc

This is called modelling error. Both Bayesian and frequentist approaches suffer from modelling error. That's what TFA talks about when mentioning the normality assumptions behind the paper's GLM. Moreover, if errors are additive, certain distributions combine together easily algebraically meaning it's easy to "marginalize" over them as a single error term. In most GLMs, there's a normally distributed error term meant to marginalize over multiple i.i.d normally distributed error terms.

> Plenty of downvotes and comments, but nothing addressing the point of the argument might suggest something.

I don't understand the point of your argument. Please clarify it.

> Here’s the experiment and here’s the data is concrete it may be bogus but it’s information. Updating probabilistic based on recursive estimates of probabilities is largely restating your assumptions.

What does this mean, concretely? Run me through an example of the problem you're bringing up. Are you saying that posterior-predictive distributions are "bogus" because they're based on prior distributions? Why? They're just based on the application of Bayes Law.

> Black swans can really throw a wrench into things

A "black swan" as Taleb states is a tail event, and this sort of analysis is definitely performed (see: https://en.wikipedia.org/wiki/Extreme_value_theory). In the case of Bayesian stats, you're specifically calculating the entire posterior distribution of the data. Tail events are visible in the tails of the posterior predictive distribution (and thus calculable) and should be able to tell you what the consequences are for a misprediction.

Re: Effect size is significantly more important than statistical significance

#74
post #56

The title's misinformation: effect-size ISN'T more important than statistical significance. The article itself makes some better points, e.g. > I worry that because of statistical ambiguity, there’s not much that can be deduced at all. , which would seem like a reasonable interpretation of the study that the article discusses. However, the title alone seems to assert a general claim about statistical interpretation t…

Not so fast. If you win your first jackpot on the first ticket. You'll require 500,000 failures (at $1 per ticket) in order to fail to reject the null hypothesis at p If you bought just ten tickets you would have a p value below 0.0000001

And that makes sense, because a p value of 0.01 says the probability of getting a sample this far from the null hypothesis is less than 1 in a million by random chance... which is what happened when you got the extremely unlikely but highly profitable answer.

edit: post was edited making this seem out of context...

Re: Effect size is significantly more important than statistical significance

#75

Earlier quoted context omitted.

The problem is that when you’re on the cusp of a new thing, unless you’re super lucky, the result will necessarily be near the noise floor. Real science is like that. But I definitely agree it’d be nice to go back and show something is true to p=.0001 or whatever. Overwhelmingly solid evidence is truly a wonderful thing, and as you say, it’s really the only way to build a solid foundation. When you engineer stuff, it…

> I’ve been thinking about this while playing Factorio: so much of our discussion and mental modeling of automation works under the assumption of perfect reliability. If you had SLIGHTLY below 100% reliability in Factorio, the game would be a terrible grind limited to small factories. So I'm making a guess here that you play with few monsters or non-aggressive monsters?

Currently playing a game to minimize pollution to try to totally avoid biter attention. Surrounded by trees, now almost entirely solar with efficiency modules.

Re: Effect size is significantly more important than statistical significance

#76
post #2

Speaking not to this study in particular necessarily, I strongly agree with the general point. Science has really been held back by an over-focusing on "significance". But I'm not really interested in a pile of hundreds of thousands of studies that establish a tiny effect with suspiciously-just-barely-significant results. I'm interested in studies that reveal robust results that are reliable enough to be built on to…

> If I were King of Science, or at least, editor of a prestigious journal, I'd want to put word out that I'm looking for papers with at least one of some sort of significant effect, or a p value of something like p = 0.0001. Yeah. That's a high bar. I know. That's the point. And study preregistration to avoid p-hacking and incentivize publishing negative results. And full availability of data, aka "open science".

I've thought about the idea of allowing people to separately publish data and analysis. Right now, data are only published if the analysis shows something interesting.

Improving the quality of measurements and data could be a rewarding pursuit, and could encourage the development of better experimental technique. And a good data set, even if it doesn't lead to an immediate result, might be useful in the future when combined with data that looks at a problem from another angle.

Granted, this is a little bit self serving: I opted out of an academic career, partially because I had no good research ideas. But I love creating experiments and generating data! Fortunately I found a niche at a company that makes measurement equipment. I deal with the quality of data, and the problem of replication, all day every day.

Re: Effect size is significantly more important than statistical significance

#77
Agree with the title, but not the contents. The study in question is actually an example of a huge effect size (10% reduction in cases just from instructing villages they should wear masks is amazing) possibly hampered by poor statistical significance (as the blog post outlines).

Re: Effect size is significantly more important than statistical significance

#78
post #71

Earlier quoted context omitted.

Nobody likes being in the control group of the first working anti-aging serum...

> Nobody likes being in the control group of the first working anti-aging serum... You only know whether it works when the study has been completed. You also only know whether the drug has (potentially) disastrous consequences when the study has been completed. Thus, I am not completely sure whether your claim holds.

You missed the working part. Success was a prerequisite to their after the fact feelings. At least some of the control group will be in old age but still alive when we know it woris. They might not know if it is infinite life (and side effects may turn it into die at 85, so some control may outlive the intervention group after the study ), but they will know on average they did worse

Re: Effect size is significantly more important than statistical significance

#79

> If most effect sizes are small or zero, then most interventions are useless. But this doesn't necessarily follow, does it? If there really were a 1.1-fold reduction in risk due to mask-wearing it could still be beneficial to encourage it. The salient issue (taking up most of the piece) seems to be not the size of the effect but rather the statistical methodology the authors employed to measure that size. The p-valu…

> If there really were a 1.1-fold reduction in risk due to mask-wearing it could still be beneficial to encourage it.

That's understating it. The study doesn't measure the reduction in risk due to mask-wearing, but rather the reduction simply from encouraging mask-wearing (which only increases actual mask wearing by a limited amount). If the study's results hold up statistically, then they're really impressive. With the caveat of course that they apply to older variants with less viral loads than Delta - it's likely Delta is more effective against masks simply due to its viral load.

> The salient issue (taking up most of the piece) seems to be not the size of the effect but rather the statistical methodology the authors employed to measure that size. The p-value isn't meaningful in the face of an incorrect model -- why isn't the answer a better model rather than just giving up?

Exactly. The irony of this article is that this is an example where effect size is actually not the issue - it's potential issues with statistical significance due to imperfect modeling, and an inability for other researchers to rerun an analysis on statistical significance, due to not publishing the raw data.

Re: Effect size is significantly more important than statistical significance

#80
post #4
post #2

Speaking not to this study in particular necessarily, I strongly agree with the general point. Science has really been held back by an over-focusing on "significance". But I'm not really interested in a pile of hundreds of thousands of studies that establish a tiny effect with suspiciously-just-barely-significant results. I'm interested in studies that reveal robust results that are reliable enough to be built on to…

> Plus, the idea that we can remove such small, noisy confounding factors is just silly. We need to look for the things that stand out from that noise floor We have found most of them, and all the easy ones. Today the interesting things are near the noise floor. 3000 years ago atoms were well below the noise floor, now we know a lot about them - most of it seems useless in daily life yet a large part of the things we…

Individual atoms, or small numbers of them, may be beneath some noise floor, but not combined atoms.

A salt crystal (Lattice of NaCl atoms) is nothing like a pure gold nugget (clump of Au atoms).

That difference is a massive effect.

So to begin with, we have this sort of massive effect which requires an explanation, which is where atoms then come in.

Maybe the right language here is not that we need an effect rather than statistical significance, but that we need a clear, unmistakable phenomenon. There has to be a phenomenon, which is then explained by research. Research cannot be inventing the phenomenon by whiffing at the faint fumes of statistical significance.

Post reply on HN