Live data from Hacker News

Effect size is significantly more important than statistical significance

argmin.net

131–140 of 168 posts

Re: Effect size is significantly more important than statistical significance

#131
post #2

Speaking not to this study in particular necessarily, I strongly agree with the general point. Science has really been held back by an over-focusing on "significance". But I'm not really interested in a pile of hundreds of thousands of studies that establish a tiny effect with suspiciously-just-barely-significant results. I'm interested in studies that reveal robust results that are reliable enough to be built on to…

> If I were King of Science, or at least, editor of a prestigious journal, I'd want to put word out that I'm looking for papers with at least one of some sort of significant effect, or a p value of something like p = 0.0001. Yeah. That's a high bar. I know. That's the point. And study preregistration to avoid p-hacking and incentivize publishing negative results. And full availability of data, aka "open science".

This paper was pretty clearly pre-specified here; https://files.givewell.org/files/DWDA%202009/IPA/Masks_RCT_P...

Re: Effect size is significantly more important than statistical significance

#132
post #13
post #4

Earlier quoted context omitted.

> Plus, the idea that we can remove such small, noisy confounding factors is just silly. We need to look for the things that stand out from that noise floor We have found most of them, and all the easy ones. Today the interesting things are near the noise floor. 3000 years ago atoms were well below the noise floor, now we know a lot about them - most of it seems useless in daily life yet a large part of the things we…

I don't think we have found most of them. I think we make it look like we've found most of them because we keep throwing money at these crap studies. Bear in mind that my criteria are two-dimensional, and I'll accept either. By all means, go back and establish your 3% effect to a p-value of 0.0001. Or 0.000000001. That makes that 3% much more interesting and useful. It'll especially be interesting and valuable when y…

So 3% is not interesting but the difference between 10^-7 and 10^-8 probability that there is no effect is interesting somehow?

Re: Effect size is significantly more important than statistical significance

#133
post #2

Speaking not to this study in particular necessarily, I strongly agree with the general point. Science has really been held back by an over-focusing on "significance". But I'm not really interested in a pile of hundreds of thousands of studies that establish a tiny effect with suspiciously-just-barely-significant results. I'm interested in studies that reveal robust results that are reliable enough to be built on to…

> or a p value of something like p = 0.0001 This has been proposed [0], albeit for a threshold of p Here's Andy Gelman and others arguing otherwise [1]. They also got like 800 scientists to sign on to the general idea of no longer using statistical significance at all [2]. [0] https://www.nature.com/articles/s41562-017-0189-z [1] http://www.stat.columbia.edu/~gelman/research/unpublished/ab... [2] https://www.nature.c…

Given the (estimated) number of scientists in the world and their general propensity to sign on to something… is 800 scientists a significant amount?

Re: Effect size is significantly more important than statistical significance

#134

Earlier quoted context omitted.

Yes, and knowing what's been tried and what has failed is important.

we tried using 0.1 mL, it didn't work we tried using 0.11 mL, it didn't work we tried using 0.12 mL, it didn't work we tried using 0.13 mL, it didn't work

    we tried using 0.10 mL, it didn't work
    we tried using 0.11 mL, it didn't work
    we tried using 0.13 mL, it didn't work
    we tried using 0.15 mL, it didn't work
    we tried using 0.17 mL, it didn't work
    we tried using 0.16 mL, it didn't work
    we tried using 0.18 mL, it didn't work
    we tried using 0.20 mL, it didn't work
    we tried using 0.14 mL, it didn't work
    we tried using 0.12 mL, it worked so we published
Do you want to know the ones that "didn't work" existed? Or are you happy with just the one that "worked" being written up in isolation?

Re: Effect size is significantly more important than statistical significance

#135
post #128

Earlier quoted context omitted.

> Prior-hacking is easier and harder to detect than p-hacking But that's comparing apples to oranges. Setting a reasonable prior is akin to frequentists interpreting the effect size (including its confidence interval) in light of deep domain knowledge. To produce a good analysis using either Bayesian or frequentist methodology (or to criticise such an analysis), you have to have deep domain knowledge. There's no gett…

> To produce a good analysis using either Bayesian or frequentist methodology (or to criticise such an analysis), you have to have deep domain knowledge. There's no getting around that, and arguably the use of p-values often lets you get away with shoddy domain knowledge. The whole problem we're facing is that it requires too much domain knowledge and detailed analysis to dismiss results that are actually just noise.…

> (you can't say anything until you've defined your prior, which requires deep domain knowledge)

Well, you can use a non-informative prior. And that's the correct choice when you genuinely don't have a better option. But you should always be able to justify that, and that in turn requires deep domain knowledge....which leads me to....

> The whole problem we're facing is that it requires too much domain knowledge and detailed analysis to dismiss results that are actually just noise.

....this is in no way a "problem" that needs fixing, by allowing shortcuts that can easily be hacked. Rather, it's a factual statement about the difficulty of drawing correct conclusions, in low Signal-to-Noise-Ratio domains. Whether you use p-values or not, and whether you use Bayesian methodology or not, you cannot get around the need to understand the data you're working with. Bad p-values are worse than none, since you have no knowledge of what error rate they actually achieve in the long-run.

> Bayesianism has no substitute for that

Yes it does. It's called Bayes factors. But as I said above, I completely disagree with your view of what a p-value is for.

Re: Effect size is significantly more important than statistical significance

#136
post #30

From the article: Ernest Rutherford is famously quoted proclaiming “If your experiment needs statistics, you ought to have done a better experiment.” “Of course, there is an existential problem arguing for large effect sizes. If most effect sizes are small or zero, then most interventions are useless. And this forces us scientists to confront our cosmic impotence, which remains a humbling and frustrating experience.”

Must be nice. Not everyone has the luxury of being able to carry out whatever experimentation they feel like. Sometimes we’re limited by what is affordable, practical, or ethical.

To take this further. Most science is a slave to grant funding. Grant funders like certain things and most of them are not biostatisticians.

That is not to say that hypercapitalism is the problem here. I think any competitive system even under socialism would have the exact same problem. Basically there are too many voices, and the ones winning are often cheating with bad statistics.

Re: Effect size is significantly more important than statistical significance

#137
post #128

Earlier quoted context omitted.

> To produce a good analysis using either Bayesian or frequentist methodology (or to criticise such an analysis), you have to have deep domain knowledge. There's no getting around that, and arguably the use of p-values often lets you get away with shoddy domain knowledge. The whole problem we're facing is that it requires too much domain knowledge and detailed analysis to dismiss results that are actually just noise.…

> (you can't say anything until you've defined your prior, which requires deep domain knowledge) Well, you can use a non-informative prior. And that's the correct choice when you genuinely don't have a better option. But you should always be able to justify that, and that in turn requires deep domain knowledge....which leads me to.... > The whole problem we're facing is that it requires too much domain knowledge and…

> Well, you can use a non-informative prior. And that's the correct choice when you genuinely don't have a better option.

At which point you've just found a more cumbersome way to do frequentist statistics. Frequentist tools aren't inconsistent with Bayes' law (they can't be, since both are valid theorems) - indeed one could say that the whole project of frequentist statistics consists of building a well-understood suite of pre-baked priors and computations that are appropriate to situations that are commonly encountered.

> ....this is in no way a "problem" that needs fixing, by allowing shortcuts that can easily be hacked. Rather, it's a factual statement about the difficulty of drawing correct conclusions, in low Signal-to-Noise-Ratio domains. Whether you use p-values or not, and whether you use Bayesian methodology or not, you cannot get around the need to understand the data you're working with.

Well, the fact is there are too many small-sample studies being produced for all or even most of them to be critically analysed by people with deep understanding. And maybe the right fix for the problem is to give the right incentives for that kind of critical analysis (e.g. by allowing that kind of analysis to count as research for the purposes of journal publications and PhD theses just as much as "the original study" does, given that a study without that kind of critical analysis cannot truly be said to represent advancing human knowledge). But if you just tell people to do Bayesian analysis instead of frequentist analysis then that's not going to magically create deep understanding - rather people will try to replace shallow frequentist analysis with shallow Bayesian analysis, and shallow Bayesian analysis is a lot less effective and more hackable.

> Yes it does. It's called Bayes factors.

But you still need a prior to compute a Bayes factor.

Re: Effect size is significantly more important than statistical significance

#138
If you have a tiny effect size on X, you probably haven't discovered a significant cause of X, but just something incidental.

For example, smoking was finally proved to cause lung cancer because the effect size was so large that the argument that 'correlation does not imply causation' became absurd: it would have required the existence of a genetic or other common cause Z that both causes people to smoke and causes them to develop cancer with correlations at least as large as between smoking and lung cancer, but there just isn't anything correlated that strongly. It would imply that almost everyone who smokes heavily does so because of Z.

Re: Effect size is significantly more important than statistical significance

#139
post #137

Earlier quoted context omitted.

> (you can't say anything until you've defined your prior, which requires deep domain knowledge) Well, you can use a non-informative prior. And that's the correct choice when you genuinely don't have a better option. But you should always be able to justify that, and that in turn requires deep domain knowledge....which leads me to.... > The whole problem we're facing is that it requires too much domain knowledge and…

> Well, you can use a non-informative prior. And that's the correct choice when you genuinely don't have a better option. At which point you've just found a more cumbersome way to do frequentist statistics. Frequentist tools aren't inconsistent with Bayes' law (they can't be, since both are valid theorems) - indeed one could say that the whole project of frequentist statistics consists of building a well-understood s…

> At which point you've just found a more cumbersome way to do frequentist statistics.

Hmm, in one way, yes...but on the other hand, Bayesian posteriors are a lot more intuitive to interpret, for most people. So I think you trade one form of convenience for another. But as you sort of hint at, the results should usually be fairly similar, whether you're doing frequentist or Bayesian analysis. So in most cases, I doubt it matters that much. Where it does matter, is when you have grounds for strong priors, that you want to take advantage of. In such cases you can improve your chances of being correct in the "here and now", if you do a Bayesian analysis. Whereas a frequentist analysis is only concerned with the asymptotic error rates. (but of course frequentist vs Bayesian is also a ladder, rather than a black and white distinction)

> Well, the fact is there are too many small-sample studies being produced for all or even most of them to be critically analysed by people with deep understanding.

And this I totally agree with. If there's one thing I dislike about academia, it's the tendency to fund low-powered studies that get nowhere. Better to go all in, with sufficient support from experienced people, in fewer and bigger studies.

Re: Effect size is significantly more important than statistical significance

#140
post #137

Earlier quoted context omitted.

> Well, you can use a non-informative prior. And that's the correct choice when you genuinely don't have a better option. At which point you've just found a more cumbersome way to do frequentist statistics. Frequentist tools aren't inconsistent with Bayes' law (they can't be, since both are valid theorems) - indeed one could say that the whole project of frequentist statistics consists of building a well-understood s…

> At which point you've just found a more cumbersome way to do frequentist statistics. Hmm, in one way, yes...but on the other hand, Bayesian posteriors are a lot more intuitive to interpret, for most people. So I think you trade one form of convenience for another. But as you sort of hint at, the results should usually be fairly similar, whether you're doing frequentist or Bayesian analysis. So in most cases, I doub…

> So in most cases, I doubt it matters that much. Where it does matter, is when you have grounds for strong priors, that you want to take advantage of. In such cases you can improve your chances of being correct in the "here and now", if you do a Bayesian analysis.

I completely agree with this - but it's exactly this dynamic that I think, at least in the current academic environment, does more harm than good. Effectively it normalizes publishing a result that's not strong enough to swamp the prior, but where you have some detailed situational argument for why a different prior should be used here. We already get every social science paper arguing that they should be allowed to use a 1-tailed t-test rather than 2-tailed because surely there's no possibility that their intervention would do more harm than good, and you need to get into the details of the paper to see why that's nonsense; letting them pick their own prior multiplies that kind of thing many times over.

Post reply on HN