Live data from Hacker News

P < 0.05 Considered Harmful

simplicityissota.substack.com

61–70 of 89 posts

Re: P < 0.05 Considered Harmful

#61
post #54

Earlier quoted context omitted.

I feel like a lot of these critiques are just straw-manning p-values consideration. Consider effect sizes - this seems to be a completely different (yes important ) question. Obviously the magnitude of the impact of the drug is important - but it isn't a replacement or "something to look at instead of p-value" because the chance that the results you saw are due to random variation is still important! You can see a ma…

I agree that you shouldn't look only at effect sizes any more than you should look only at p-values. (What I would actually prefer you to do, where you can figure out a good way to do it, is to compute a posterior probability distribution and look at the whole distribution. Then you can look at its mean or median or mode or something to get a point estimate of effect size, you can look at how much of the distribution…

What I have discovered after working in medicine for pretty long is that many biologists and MDs think p-values are a measure of effect sizes. Even a reviewer from Nature thought that, which is incredibly disturbing.

p-values were created to facilitate rigorous inference with minimal computation, which was the norm during the first half of the 20th century. For those who work on a frequentist framework, inference should be done using a likelihood-based approach plus model selection, e.g. AIC. It's makes it much harder to lie.

Re: P < 0.05 Considered Harmful

#62

Still not significant: https://mchankins.wordpress.com/2013/04/21/still-not-signifi... Nothing like spending 10 minutes reading a paper to see results which are likely nonsense. However, it pales in comparison to spending 3 weeks trying to replicate popular works... only to find it doesn't generalize... you know that ROC was likely from cooked data-sets confounded with systematic compression artifact errors... likely…

What a bananas list! Thanks for sharing it.

Re: P < 0.05 Considered Harmful

#63
post #32

This post is amazing. When I see a post "P But no, this one isn't from a frequentist, or from a Bayesian. It's from a techbro whose solution isn't any kind of multiple hypothesis correction or getting Bayes-pilled, it's to say "Why not just admit you want something akin to p = 0.25 in the first place?" for 'ship criteria' in the only stats context he appears to know which is A/B testing, talking about Maslow hierarch…

> "Why not just admit you want something akin to p = 0.25 in the first place?" for 'ship criteria' That's culture shock for me - I guess this is why I don't work at startups.

p=0.2 doesn't work too well for ship criteria for medicine.

p=0.2 for "this reordering of the landing text improves the rate conversion events" is fine. People make changes based on less information all the time. Waiting for certainty has its own expenses.

Re: P < 0.05 Considered Harmful

#64
post #61
post #54

Earlier quoted context omitted.

I agree that you shouldn't look only at effect sizes any more than you should look only at p-values. (What I would actually prefer you to do, where you can figure out a good way to do it, is to compute a posterior probability distribution and look at the whole distribution. Then you can look at its mean or median or mode or something to get a point estimate of effect size, you can look at how much of the distribution…

What I have discovered after working in medicine for pretty long is that many biologists and MDs think p-values are a measure of effect sizes. Even a reviewer from Nature thought that, which is incredibly disturbing. p-values were created to facilitate rigorous inference with minimal computation, which was the norm during the first half of the 20th century. For those who work on a frequentist framework, inference sho…

AIC is an estimate of prediction error. I would caution against using it for selecting a model for the purpose of inference of e.g. population parameters from some dataset (without producing some additional justification that this is a sensible thing to do). Also, uncertainty quantification after data-dependent model selection can be tricky.

Best practice (as I understand it) is to fix the model ahead of time, before seeing the data, if possible (as in a randomized controlled trial of a new medicine, etc.).

Re: P < 0.05 Considered Harmful

#65
The fundamental error that causes misuse of p-values (and statistics in general) is misunderstanding what statistics is in the first place. Statistics is applied epistemology. There is no algorithm that fits all situations where you're trying to think about and learn new things. It's just hard and we have to deal with that.

Arguably, the main point of p-values is trying to prevent people who really know better from "cheating" by reporting "interesting" results from very low sample sizes. Having a very rigid framework as a rejection criteria helps with this. However, the scientific community and system are not capable of dealing with "real cheating" that includes fabrication of data. Also, any such rigid metric is going to be gamed. Of course, there are some people who would also cheat themselves and maybe learning about p-values makes this less frequent. But such people using them don't understand what those values are telling them, beyond "yes, I can publish this". This is counterproductive because it prevents deep thinking and thorough investigation.

Most scientists realize that in practice there is no simple list of objective criteria that tells you what experiments to perform and how exactly to interpret the results. This takes a lot of work, careful thinking, trying different things and very much benefits from collaboration. But who's got time for that? There's papers to publish and grant applications to write. So p-values it is. Or maybe some other thing eventually, that also won't solve the fundamental problem.

Re: P < 0.05 Considered Harmful

#66

So much has been said about p-values and null hypothesis significance testing (NHST) that one blog post probably won't change anybody's opinion. But I want to recommend the wonderful paper "the null ritual" [1] by Gerd Gigerenzer et al. It shows that a precise understanding of what a p-value means is extremely rare even among statistics lecturers, and, more interestingly, that there have always been fundamentally dif…

Thanks for this, looking forward to read it! I would also recommend in turn the paper by Greenland et al (2016) called "Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations.". It's a great read.

I recently had a chat with a colleague who has "abandoned p-values for bayes factors" in their research, on how, in principle, there's nothing stopping you from having a "bayesian" p-value (i.e. the definition of the p-value at its most general can easily accommodate bayesian inference, priors, posteriors, etc). The counter-retort was more or less "no it can't, educate yourself, bayes factors are better" and didn't want to hear about it. It made me sad.

p-values are an incredibly insightful device. But because most people (ab)use it in the same way most people abuse normality assumptions or ordinal scales as continuous, it's gotten a bad rep and means something entirely different to most now by default.

Re: P < 0.05 Considered Harmful

#67
post #60
post #18

Earlier quoted context omitted.

The "stasis" and "arbitrarily adjustments" regimes that I wrote about are certainly ones that I've seen, which don't rely solely on p But let me turn that around and ask: what product decision regime do you see most often or think would be the most relevant to use as an example? I'd be happy to hear your perspective and make sure I keep it in mind for future blogs.

Hi I wrote some other response in this thread where I called you a techbro, so sorry about that I probably count as one too I wasn't trying to insult you too much. Anyway when you say "Furthermore, it's not only about whether 0.05 is the sole criteria, but also about whether it's a useful criteria for us to highlight at all, depending on whether the anchoring effect of it is damaging relative to alternatives." I love…

I've been called worse things, and at this stage in life it is hard to be offended by people who don't know me well enough to give an insightful insult. But I hadn't responded to the earlier comment because it was more generally antagonistic and seemed to reflect a (possibly intentional?) misinterpreted and negative reading of the blog. "Don't feed the trolls", as the saying goes.

I started writing that particular blog with different experimental methods in mind, but wrote so much as a prerequisite that I wanted to stop and make that first part a standalone post. My last paragraph was supposed to make it clear that this was a launching off point, rather than a summation of everything I know about decision theory or experimentation.

Thanks for writing more substance into this comment. These opinions are sensible, I agree with much that you wrote here. On one: I do use MABs and see them as a feasible method for many organizations, and at some point I'd like to write about some challenges with those too.

Wishing you the best in your techbro or ftxbro or bro life too!

Re: P < 0.05 Considered Harmful

#68
post #32

This post is amazing. When I see a post "P But no, this one isn't from a frequentist, or from a Bayesian. It's from a techbro whose solution isn't any kind of multiple hypothesis correction or getting Bayes-pilled, it's to say "Why not just admit you want something akin to p = 0.25 in the first place?" for 'ship criteria' in the only stats context he appears to know which is A/B testing, talking about Maslow hierarch…

[flagged]

Re: P < 0.05 Considered Harmful

#69
post #15

The article is all about why "0.05" might be a bad value to choose. But, more fundamentally, p is often the wrong thing to be looking at in the first place. 1. Effect sizes. Suppose you are a doctor or a patient and you are interested in two drugs. Both are known to be safe (maybe they've been used for decades for some problem other than the one you're now facing). As for efficacy against the problem you have, one ha…

It’s worth pointing out that processes like meta analyses ought to be able to catch #2. The problem is, GIGO. A lot of poorly designed studies (no real control group, failure to control for other relevant variables, other methodological errors) make meta analysis outcomes unreliable.

An area I follow closely is exercise science and I am amazed that researchers are able to get grants for some of the research they do. For instance, sometimes researchers aim to compare, say, the amount of hypertrophy on one training program versus another. They’ll use a study or maybe 15 individuals. The group on intervention A will experience 50% higher hypertrophy than intervention B, at p=0.08, and they’ll conclude that there’s no difference in hypertrophy between the two protocols rather than suggesting an increase in statistical power.

Another great example is studies whose interventions fail to produce any muscle growth in new trainees. They’ll compare two programs, one with, say, higher training volume, and one with lower training volume. Both fail to produce results for some reason. They conclude that variable is not important, rather than perhaps concluding that their interventions are poorly designed since a population that is begging to gain muscle couldn’t gain anything from it.

Re: P < 0.05 Considered Harmful

#70
post #15

The article is all about why "0.05" might be a bad value to choose. But, more fundamentally, p is often the wrong thing to be looking at in the first place. 1. Effect sizes. Suppose you are a doctor or a patient and you are interested in two drugs. Both are known to be safe (maybe they've been used for decades for some problem other than the one you're now facing). As for efficacy against the problem you have, one ha…

> The article is all about why "0.05" might be a bad value to choose.

No. It's more about why choosing a value to serve as the default choice is a bad idea in the first place. The specific value chosen as the default itself (i.e., 0.05 in this case) is irrelevant.

The idea is that the value you choose should reflect some prior knowledge about the problem. Therefore choosing 5% all the time would be somewhat analogous to a bayesian choosing a specific default gaussian as their prior: it defeats the point of choosing a prior, and it's actively harmful if it sabotages what you should actually be using as a prior, because it's not even uninformative, it's a highly opinionated prior instead.

As for points 1 to 3, technically I agree, but there's a lot of misdirection involved.

The point on effect sizes is true (and indeed something many people get wrong), but it is contrived to make effect sizes more useful than p-values. In which case, the obvious answer is, you should be choosing an example where reporting a p-value is more important than an effect size. One way to look at a p-value is as a ranking, which would be useful for comparing between effect sizes of incomparable units. Is a student with a 19/20 grade from a european school better than an american student with a 4.6 GPA? Reducing the compatibility scores to rankings, can help you compare these two effect sizes immediately.

Prior probability, similarly. "If" you interpret things as you did, then yes, p-values suck. But you're not supposed to. What the p-value tells you is "the data and model are 'this much' compatible". It's up to you then to say "and therefore" vs "must have been an atypical sample". In other words, there is still space for a prior here. And in theory, you are free to repurpose the compatibility score of a p-value to introduce this prior directly (though nobody does this in practice).

Regarding p=0.05 vs p=0.001 not mattering; of course they do. But only if they're used as compatibility rankings as opposed to decision thresholds. If you compare two models, and one has p=0.05 and the other has p=0.001, this tells you two things: a) they are both very incompatible with the data, b) the latter is a lot more incompatible than the former. The problem is not that people use p-values, the problem that people abuse them, to make decisions that are not necessarily crisply supported by the p-values used to push them. But this could be said of any metric. I have actively seen people propose "deciding in favour of a model if the Bayes Factor is > 3". This is exactly the same faulty logic, and the fact that BF is somehow "bayesian" won't protect subsequent researchers who use this heuristic from entering a new reproducibility crisis.

Post reply on HN