Live data from Hacker News

P < 0.05 Considered Harmful

simplicityissota.substack.com

71–80 of 89 posts

Re: P < 0.05 Considered Harmful

#71
post #58

Earlier quoted context omitted.

You can always encode the prior. If you take the frequentist approach and ignore bayesian concepts, it’s the same as just going bayesian but with an “uninformative prior” (constant distribution). The only question is… would you rather be up front and explicit about your assumptions, or not? An uninformative prior is an assumption, even if it’s the one that doesn’t bias the posterior (note that here “bias” is not a ba…

There is potentially bias (of the bad word variant) introduced by the mismatch between the prior in your own mind and the distribution and params you choose to try to approximate that, especially if you're trying to pick out a distribution with a nice posterior conjugate. I'm also not sure why everyone perceived my comment as anti-bayesian.

I did not perceive your comment as anti-bayesian, or at least not necessarily so! :)

But are you sure you know what I meant when I said “uninformative prior”? Because choosing an uninformative prior does not involve choosing any parameters: there is only one uninformative prior, and it’s the constant (flat) distribution which assigns equal probability to every value. It encodes no information and does not bias the posterior or result. It is the one and only mathematically-neutral prior. You can think of it as being a bit like an “identity function”.

Re: P < 0.05 Considered Harmful

#72
post #61

Earlier quoted context omitted.

What I have discovered after working in medicine for pretty long is that many biologists and MDs think p-values are a measure of effect sizes. Even a reviewer from Nature thought that, which is incredibly disturbing. p-values were created to facilitate rigorous inference with minimal computation, which was the norm during the first half of the 20th century. For those who work on a frequentist framework, inference sho…

AIC is an estimate of prediction error. I would caution against using it for selecting a model for the purpose of inference of e.g. population parameters from some dataset (without producing some additional justification that this is a sensible thing to do). Also, uncertainty quantification after data-dependent model selection can be tricky. Best practice (as I understand it) is to fix the model ahead of time, before…

And it is not uncommon that an intentionally bad model (low AIC) will be used for inference on a parameter when one wants to test the robustness of the parameter to covariates.

Re: P < 0.05 Considered Harmful

#73
post #15

The article is all about why "0.05" might be a bad value to choose. But, more fundamentally, p is often the wrong thing to be looking at in the first place. 1. Effect sizes. Suppose you are a doctor or a patient and you are interested in two drugs. Both are known to be safe (maybe they've been used for decades for some problem other than the one you're now facing). As for efficacy against the problem you have, one ha…

I feel like a lot of these critiques are just straw-manning p-values consideration. Consider effect sizes - this seems to be a completely different (yes important ) question. Obviously the magnitude of the impact of the drug is important - but it isn't a replacement or "something to look at instead of p-value" because the chance that the results you saw are due to random variation is still important! You can see a ma…

Indeed.

Ive worked with clinical trials and nobody says “oh, p Doctors do a much more sophisticated deep dive into the data and statistical significance is maybe 1 of a 10 factors they consider.

And hell, even if it is statistically significant, doctors may not use a drug for many other reasons.

It seems like the people who criticize p-values aren't that familiar with how they are actually used.

Re: P < 0.05 Considered Harmful

#75

Earlier quoted context omitted.

Your points are theoretically correct, and probably the reason why many statisticians still regard p-values and HNST favorably. But looking at the practical application, in particular the replication crisis, specification curve analysis, de facto power of published studies and many more, we see that there is an immense practical problem and p-values are not making it better. We need to criticize p-values and NHST har…

The items you listed are certainly problems, but p-values don't have much to do with them, as far as I can see. Poor power is an experimental design problem, not a problem with the analysis technique. Not reporting all analyses is a data censoring problem (this is what I understand "specification curve analysis" to mean, based on some Googling - let me know if I misinterpreted). Again, this can't really be fixed at t…

I think part of the problem with p-values and NHST is that it encourages (or doesn't discourage) underpowered studies. That's because p-hacking benefits from the noise of underpowered studies. If you can test a large number of models and only report the significant one then an underpowered study with high type I error rate gives you a greater chance of a significant result.

So I think you are correct that properly powering studies is the crucial thing, but the incentives are against fixing this as long as lone p-values are publishable.

Re: P < 0.05 Considered Harmful

#76
post #50

Earlier quoted context omitted.

Some games (both computer games and physical board games) intentionally use "shuffled randomness" where e.g. for percentile fail/success rolls you'd take numbers from 1-100 and use true randomness to shuffle that list; in this way the overall probability is the same, but has a substantially different feel as it's impossible for someone to have bad/good luck throughout the whole game and things like "gambler's fallacy…

> the overall probability is the same Only for the very first roll. After that, the outcome becomes more and more predictable for each roll.

Just as with any standard PRNG (e.g. Mersenne twister) the outcome may be fully predictable with some information about the state or the previous rolls, but its still usable; and the general distribution of values matches what is desired for that section of the game.

Re: P < 0.05 Considered Harmful

#77

Earlier quoted context omitted.

The items you listed are certainly problems, but p-values don't have much to do with them, as far as I can see. Poor power is an experimental design problem, not a problem with the analysis technique. Not reporting all analyses is a data censoring problem (this is what I understand "specification curve analysis" to mean, based on some Googling - let me know if I misinterpreted). Again, this can't really be fixed at t…

I think part of the problem with p-values and NHST is that it encourages (or doesn't discourage) underpowered studies. That's because p-hacking benefits from the noise of underpowered studies. If you can test a large number of models and only report the significant one then an underpowered study with high type I error rate gives you a greater chance of a significant result. So I think you are correct that properly po…

But here the issue is the uncorrected multiple testing and under-reporting of results, not the p-values themselves. Any criterion for judging the presence of an effect is going to suffer from the same issue, if researchers don't pre-register and report all of their analyses (since otherwise you have censored data, "researcher degrees of freedom," and so on). This is really a a problem with the design and reporting of studies, not the analysis method.

Re: P < 0.05 Considered Harmful

#78
> Neutral changes aren’t costly on our users, so while we should be somewhat averse to wasting time and adding tech debt for neutral changes, it isn’t the end of the world.

While I think most of the points in the post are reasonable, I strongly disagree with the idea that neutral changes aren't costly. Some neutral changes are good because they are a part of a larger strategy or vision, but in my experience most neutral features are simply change for changes sake. Any reasonablely large company that ships all/most neutral tests is going to end up feeling like a bloated, unfocused mess.

Re: P < 0.05 Considered Harmful

#79
post #32

This post is amazing. When I see a post "P But no, this one isn't from a frequentist, or from a Bayesian. It's from a techbro whose solution isn't any kind of multiple hypothesis correction or getting Bayes-pilled, it's to say "Why not just admit you want something akin to p = 0.25 in the first place?" for 'ship criteria' in the only stats context he appears to know which is A/B testing, talking about Maslow hierarch…

It made the front page because of the title. Bashing p-values is a formula for upvotes, regardless of the content.

A lot of people associate p-values with p-hacking and replication crisis, which they read about on HN a few times while sitting on the can. Therefore, in their expert opinions, p-values are bad.

Re: P < 0.05 Considered Harmful

#80
post #50

Earlier quoted context omitted.

> the overall probability is the same Only for the very first roll. After that, the outcome becomes more and more predictable for each roll.

Just as with any standard PRNG (e.g. Mersenne twister) the outcome may be fully predictable with some information about the state or the previous rolls, but its still usable; and the general distribution of values matches what is desired for that section of the game.

That’s not what I meant. Say that the first roll was a 1. Then you know that for at least 100 rolls, you won’t get a 1 anymore. And with more rolls, more numbers are eliminated.
Post reply on HN