Live data from Hacker News

P < 0.05 Considered Harmful

simplicityissota.substack.com

81–89 of 89 posts

Re: P < 0.05 Considered Harmful

#81
post #15

The article is all about why "0.05" might be a bad value to choose. But, more fundamentally, p is often the wrong thing to be looking at in the first place. 1. Effect sizes. Suppose you are a doctor or a patient and you are interested in two drugs. Both are known to be safe (maybe they've been used for decades for some problem other than the one you're now facing). As for efficacy against the problem you have, one ha…

I find it interesting that the prior probability example always happens to be the one where the prior is correct, while the opposite direction is never given more than half a line. You can very easily turn it around so that an unreasonably large amount of evidence is needed to move a strongly believed wrong prior.

My favourite example is to do that with normal distributions and the posterior ends up as "let's meet in the middle and agree that we both were wrong" -a very low variance posterior with mean right in the middle between prior and data.

Re: P < 0.05 Considered Harmful

#82
post #58

Earlier quoted context omitted.

I feel like a lot of these critiques are just straw-manning p-values consideration. Consider effect sizes - this seems to be a completely different (yes important ) question. Obviously the magnitude of the impact of the drug is important - but it isn't a replacement or "something to look at instead of p-value" because the chance that the results you saw are due to random variation is still important! You can see a ma…

You can always encode the prior. If you take the frequentist approach and ignore bayesian concepts, it’s the same as just going bayesian but with an “uninformative prior” (constant distribution). The only question is… would you rather be up front and explicit about your assumptions, or not? An uninformative prior is an assumption, even if it’s the one that doesn’t bias the posterior (note that here “bias” is not a ba…

Depending on the hypothesis space what you call the "uninformative prior" does not exist in the frequentist approach. If you search for a real value, then the uninformative prior is a uniform distribution on the infinite line. This distribution does not normalize and is off-limits to bayesians.

Ultimately, I think you are strawmanning frequentism here. Just because the log likelihood is sometimes the same as the map does not imply that they have the same meaning. This is why computed uncertainties of both approaches are often not the same and have a not-so-subtle difference in their interpretation. The one computes uncertainty in belief, the other imprecision of an experiment. You can't summarize that with "do you want to be explicit about assumptions".

Re: P < 0.05 Considered Harmful

#83
post #63

Earlier quoted context omitted.

> "Why not just admit you want something akin to p = 0.25 in the first place?" for 'ship criteria' That's culture shock for me - I guess this is why I don't work at startups.

p=0.2 doesn't work too well for ship criteria for medicine. p=0.2 for "this reordering of the landing text improves the rate conversion events" is fine. People make changes based on less information all the time. Waiting for certainty has its own expenses.

Not that fine. The author claims that neutral changes are not costly for users so the only reason to avoid them is to avoid wasting time/money. But that's not really true. Pointless churn annoys users. If your P threshold isn't low enough then you can get stuck in an endless treadmill of making changes that you thought would have benefit but which don't actually do anything because your threshold for something being considered significant is too loose.

Re: P < 0.05 Considered Harmful

#84
post #82
post #58

Earlier quoted context omitted.

You can always encode the prior. If you take the frequentist approach and ignore bayesian concepts, it’s the same as just going bayesian but with an “uninformative prior” (constant distribution). The only question is… would you rather be up front and explicit about your assumptions, or not? An uninformative prior is an assumption, even if it’s the one that doesn’t bias the posterior (note that here “bias” is not a ba…

Depending on the hypothesis space what you call the "uninformative prior" does not exist in the frequentist approach. If you search for a real value, then the uninformative prior is a uniform distribution on the infinite line. This distribution does not normalize and is off-limits to bayesians. Ultimately, I think you are strawmanning frequentism here. Just because the log likelihood is sometimes the same as the map…

Nobody normalizes an uninformative prior on its own. Normalization only happens when you get your posterior.

I am not straw-manning anything: bayesian methods are a generalization of frequentist methods. The equality / isomorphism to frequentist methods, in the special case of the uninformative prior, is commonly demonstrated in introductory (undergrad-level) bayesian textbooks and is in fact trivial. One need not even talk of any infinities: if you’re doing discrete observations, infinities never show up (and in that case the prior normalizes just fine). And if you’re curve-fitting (ie.: using parametric methods), the “infinite” line goes away as soon as you multiply your uninformative prior with your likelihood.

Re: P < 0.05 Considered Harmful

#85
post #15

The article is all about why "0.05" might be a bad value to choose. But, more fundamentally, p is often the wrong thing to be looking at in the first place. 1. Effect sizes. Suppose you are a doctor or a patient and you are interested in two drugs. Both are known to be safe (maybe they've been used for decades for some problem other than the one you're now facing). As for efficacy against the problem you have, one ha…

> The article is all about why "0.05" might be a bad value to choose. No. It's more about why choosing a value to serve as the default choice is a bad idea in the first place. The specific value chosen as the default itself (i.e., 0.05 in this case) is irrelevant. The idea is that the value you choose should reflect some prior knowledge about the problem. Therefore choosing 5% all the time would be somewhat analogous…

I agree that the article is somewhat about the idea that having a Standard Default p-value Threshold is unwise, but I think it's mostly suggesting that in its particular context it's generally better to use a larger p-value. "Defaults matter", the author says (not "Defaults are a trap"). "Do we need that much risk aversion?" (not "Sometimes we need more risk aversion, sometimes less"). "If we were starting over ... I doubt that's what we would pick" (not "If we were starting over, I think we'd have a process that begins by considering what p-value threshold would be appropriate"). At the end, he considers situations where the criterion used is strict and "the company progresses too slowly" (but not ones where it's lax and "the company moves too fast and breaks too many things" or "the company thrashes about unstably") and "Why not just admit you want something akin to p=0.25 in the first place" (not "... something akin to p=0.25 some of the time, and something akin to p=0.001 some of the time?") and "I believe that nobody wants to write down a large p-value threshold, because it feels unscientific" (and nothing about reasons why people might be reluctant to write down an unusually small p-value threshold).

I agree, in case it needs to be said, that if you are doing hypothesis testing then you should adapt your choice of thresholds to your (explicit or implicit) prior. That was approximately the second of the three points I made.

"You should be choosing an example where reporting a p-value is more important than an effect size." Why should I be doing that? I think that almost always both are important (more precisely, I think that almost always thinking of what you've learned from an experiment into a p-value and a point-estimate effect size is suboptimal, but you want the information both of those things are trying to tell you). And if I've correctly understood what you say about rankings, I completely disagree with it; almost none of the time is the p-value what you should be ranking things on. Again, consider those two medications in my hypothetical example: I do not believe any reasonable person would think that the right way to rank them relative to one another matches the p-values. (For comparing two individual students? Yeah, maybe, kinda. But approximately 0% of cases where people compute p-values are analogous to that.)

I agree that picking any single metric and applying brainless black-and-white rules using it is liable to get you in trouble, whatever that metric is. But some metrics are better than others, and a fixed p-value threshold (1) has been a widespread common practice and (2) is for most purposes a really bad idea.

Re: P < 0.05 Considered Harmful

#86
post #78

> Neutral changes aren’t costly on our users, so while we should be somewhat averse to wasting time and adding tech debt for neutral changes, it isn’t the end of the world. While I think most of the points in the post are reasonable, I strongly disagree with the idea that neutral changes aren't costly. Some neutral changes are good because they are a part of a larger strategy or vision, but in my experience most neut…

Does the sentence imply that they aren't costly? Would you write "strongly averse" instead of "somewhat averse" or "it is the end of the world" instead of "it isn't the end of the world"? If so, what language would you use to convey that strongly negative changes are that much worse than neutral ones?

Re: P < 0.05 Considered Harmful

#87
post #63

Earlier quoted context omitted.

p=0.2 doesn't work too well for ship criteria for medicine. p=0.2 for "this reordering of the landing text improves the rate conversion events" is fine. People make changes based on less information all the time. Waiting for certainty has its own expenses.

Not that fine. The author claims that neutral changes are not costly for users so the only reason to avoid them is to avoid wasting time/money. But that's not really true. Pointless churn annoys users. If your P threshold isn't low enough then you can get stuck in an endless treadmill of making changes that you thought would have benefit but which don't actually do anything because your threshold for something being…

> Pointless churn annoys users.

Sure, excessive rate of change is bad. Even if it's all validated and well-proven.

> If your P threshold isn't low enough then you can get stuck in an endless treadmill of making changes

To measure P, I have to make the change.

There's plenty of situations where accumulating enough data to meet a p<0.05 threshold would take months. If my prior is that it's a reasonably good change and has a decent chance to be helpful, and then I take a measurement that makes that roughly 5x more likely... it's OK to push the button. There are many decisions in business that have to be made with much less information or statistical proof than this.

Re: P < 0.05 Considered Harmful

#89
post #85

Earlier quoted context omitted.

> The article is all about why "0.05" might be a bad value to choose. No. It's more about why choosing a value to serve as the default choice is a bad idea in the first place. The specific value chosen as the default itself (i.e., 0.05 in this case) is irrelevant. The idea is that the value you choose should reflect some prior knowledge about the problem. Therefore choosing 5% all the time would be somewhat analogous…

I agree that the article is somewhat about the idea that having a Standard Default p-value Threshold is unwise, but I think it's mostly suggesting that in its particular context it's generally better to use a larger p-value. "Defaults matter", the author says (not "Defaults are a trap"). "Do we need that much risk aversion?" (not "Sometimes we need more risk aversion, sometimes less"). "If we were starting over ... I…

Yes; as I said, wasn't disagreeing per se, just clarifying important context more than anything.

> I completely disagree with it; almost none of the time is the p-value what you should be ranking things on.

Not necessarily true; people compare models on the basis of their p-values all the time; but in any case, that's not what I was referring to here. I was saying the p-value in itself expresses a ranking. That's just what a p-value is by definition: a ranking of the particular compatibility score of the data-event under consideration, with respect to the whole domain of compatibility scores resulting from all possible data-events considered.

And yes, when the score itself is more important than the ranking, it's prudent to focus on that; when the p-value "ranking" is more important, it's prudent to focus on that instead; ideally, people should consider both.

Having said that, the medical example is a weird one. There's a specific sentence there that I have a problem with:

> It's more uncertain because the sample size is smaller [...] I would definitely prefer the second drug.

Perhaps this is simply a misnomer, since one can only be so precise when using language, but as expressed above, I feel this is wrong, as it alludes to an expression of posterior probability, which is not information given to you by a p-value in this setting. More specifically you should be saying it's not "precise". The subtle difference being, you are interpreting this as a "despite the wide range, it's most likely to be around 0.1, so I'll take my chances", whereas what the confidence interval is really telling you is "the data would be atypical for models outside this large range, but not inside", meaning there is an almost equal chance of the drug harming you as there is giving you a benefit. So, no, I wouldn't prefer the 2nd drug; I'd actually prefer the 1st one which has a guaranteed, albeit smaller benefit. This is a very different statement from a credible interval which more or less would state that 0.1 is the most likely value.

So in fact, the p-value would be a better measure of comparison here as well. You're comparing the compatibility of two different data-events (which in this case reflects samplings) against the same (null) hypothesis; so you should prefer the event that is least compatible with the null (and ideally, you should propose an alternative hypothesis, and confirm that it demonstrates higher compatibility with that one as well). But having said this, comparing p-values typically comes in the reverse scenario, were one has a single data-event, and wants to compare two different hypotheses against the same data-event, since that gives information on which model is most compatible with what you observed.

Anyway. Sorry, I know you know, I'm just enjoying the exposition :)

Post reply on HN