Live data from Hacker News

How to avoid P hacking

nature.com

61–70 of 91 posts

Re: How to avoid P hacking

#63
Reading this article tbh causes second hand embarrassment. Ostensibly it’s targeting professional scientists using the brand of a prestigious journal, yet it has a vibe of explaining ethics and common sense to school kids. We’ve come to the point of having to explain to PhDs why cherry picking data is bad.

I’m not criticizing the article, rather bemoaning the fact that it’s needed. Of course the problem is not just with the much maligned social sciences, it’s physics and computer science too. The controversy around Microsoft’s topological qubits, a super complex topic, in part involved the most basic kind of this nonsense, something like including 4 samples of 20 measured in the paper iirc.

The community needs to get its shit together. The world we’re living in now, the post truth era, is the result of many factors but this is one of them. The loss of faith in science is partially a self-inflicted wound.

Re: How to avoid P hacking

#64

Like the old saying goes, "It is difficult to get a researcher to stop P hacking, when his career depends on his not stopping P hacking."

It is an old saying, and I’m not sure there’s much use to it as it feels like a mitigation.

No doubt the system needs to change, but lots of careers benefit from cheating or unethical behavior. It doesn’t rationalize it or force a choice on anyone.

Re: How to avoid P hacking

#65

I was heavily encouraged to do what would later be called “p-hacking”, but it looked different from what they describe here. This article describes p-hacks for people that aren’t into math/stats. I always ended up p hacking because I was into stats methods. Somebody would say “here’s an old dataset that didn’t work out, I bet you can use one of those new stats methods you’re always reading about to find a cool effect…

The practice you describe is called data dredging though. The thing about it is that you do not know enough experimental design details to make sure it was all on the up, especially worse the older the dataset gets. Normally when doing that you need a multiple comparison corrections and conservative stats. That won't get you published though, or if you do get published you won't get noticed except by someone running…

If you "dredge" any data set (even the one you can 100% trust) over and over with random hypotheses until p-value is <0.05, you will eventually (actually, pretty quickly) support some false hypothesis. That's why "data dredging" is also p-hacking.

Re: How to avoid P hacking

#66
post #62
post #61

>If you need statistics, you did the wrong experiment. ~Ernest Rutherford.

>If you don't need statistics, you did the wrong experiment. ~Psychologists >What are statistics? ~Computer scientists

Psychologists are notoriously bad at statistics though

Re: How to avoid P hacking

#67

Earlier quoted context omitted.

The practice you describe is called data dredging though. The thing about it is that you do not know enough experimental design details to make sure it was all on the up, especially worse the older the dataset gets. Normally when doing that you need a multiple comparison corrections and conservative stats. That won't get you published though, or if you do get published you won't get noticed except by someone running…

If you "dredge" any data set (even the one you can 100% trust) over and over with random hypotheses until p-value is <0.05, you will eventually (actually, pretty quickly) support some false hypothesis. That's why "data dredging" is also p-hacking.

Yes, as I understand it there is bias inherent in any dataset due to the fact it is a sample. Data dredging is just looking for that bias. You could do that, but then you'd have to confirm with a new experiment.

Re: How to avoid P hacking

#68

I was heavily encouraged to do what would later be called “p-hacking”, but it looked different from what they describe here. This article describes p-hacks for people that aren’t into math/stats. I always ended up p hacking because I was into stats methods. Somebody would say “here’s an old dataset that didn’t work out, I bet you can use one of those new stats methods you’re always reading about to find a cool effect…

As long as there is transparency about the process, I think this sort of thing is basically fine. It's roughly at the level of observational science rather than experimental science, and it can help lead to new research to validate the effect discovered.

Where this gets dangerous is when it is taken at face value, either in scientific circles, or, more common, journalistic circles.

Re: How to avoid P hacking

#69
The article cuts off for me so I do not know if they talk about this, but preregistration has to be part of the conversation moving forward.

And it has to have teeth -- withdrawn studies have to have a reputational risk that affects the credibility of future studies, even if it means publishing a retrospective or a null result in a minor journal.

Re: How to avoid P hacking

#70

> Stopping an experiment once you find a significant effect but before you reach your predetermined sample size is classic P hacking. Although much of the article is basic common sense, and although I'm not a statistician, I had to seriously question the author's understanding of statistics at this point. The predetermined sample size (statistical power) is usually based on an assumption made about the effect size; i…

No, it's generally not valid -- it will depend on the specifics of the test (especially if the test is valid only asymptotically). You need some method that supports sequential inference. Nowadays your best bet is probably some sort of anytime-valid method from the e-value literature https://en.wikipedia.org/wiki/E-values https://projecteuclid.org/journals/statistical-science/volum...
Post reply on HN