~Ernest Rutherford.
How to avoid P hacking
61–70 of 91 posts
Re: How to avoid P hacking
#62>If you need statistics, you did the wrong experiment. ~Ernest Rutherford.
~Psychologists
>What are statistics?
~Computer scientists
Re: How to avoid P hacking
#63I’m not criticizing the article, rather bemoaning the fact that it’s needed. Of course the problem is not just with the much maligned social sciences, it’s physics and computer science too. The controversy around Microsoft’s topological qubits, a super complex topic, in part involved the most basic kind of this nonsense, something like including 4 samples of 20 measured in the paper iirc.
The community needs to get its shit together. The world we’re living in now, the post truth era, is the result of many factors but this is one of them. The loss of faith in science is partially a self-inflicted wound.
Re: How to avoid P hacking
#64Like the old saying goes, "It is difficult to get a researcher to stop P hacking, when his career depends on his not stopping P hacking."
No doubt the system needs to change, but lots of careers benefit from cheating or unethical behavior. It doesn’t rationalize it or force a choice on anyone.
Re: How to avoid P hacking
#65I was heavily encouraged to do what would later be called “p-hacking”, but it looked different from what they describe here. This article describes p-hacks for people that aren’t into math/stats. I always ended up p hacking because I was into stats methods. Somebody would say “here’s an old dataset that didn’t work out, I bet you can use one of those new stats methods you’re always reading about to find a cool effect…
The practice you describe is called data dredging though. The thing about it is that you do not know enough experimental design details to make sure it was all on the up, especially worse the older the dataset gets. Normally when doing that you need a multiple comparison corrections and conservative stats. That won't get you published though, or if you do get published you won't get noticed except by someone running…
Re: How to avoid P hacking
#66Re: How to avoid P hacking
#67Earlier quoted context omitted.
The practice you describe is called data dredging though. The thing about it is that you do not know enough experimental design details to make sure it was all on the up, especially worse the older the dataset gets. Normally when doing that you need a multiple comparison corrections and conservative stats. That won't get you published though, or if you do get published you won't get noticed except by someone running…
If you "dredge" any data set (even the one you can 100% trust) over and over with random hypotheses until p-value is <0.05, you will eventually (actually, pretty quickly) support some false hypothesis. That's why "data dredging" is also p-hacking.
Re: How to avoid P hacking
#68I was heavily encouraged to do what would later be called “p-hacking”, but it looked different from what they describe here. This article describes p-hacks for people that aren’t into math/stats. I always ended up p hacking because I was into stats methods. Somebody would say “here’s an old dataset that didn’t work out, I bet you can use one of those new stats methods you’re always reading about to find a cool effect…
Where this gets dangerous is when it is taken at face value, either in scientific circles, or, more common, journalistic circles.
Re: How to avoid P hacking
#69And it has to have teeth -- withdrawn studies have to have a reputational risk that affects the credibility of future studies, even if it means publishing a retrospective or a null result in a minor journal.
Re: How to avoid P hacking
#70> Stopping an experiment once you find a significant effect but before you reach your predetermined sample size is classic P hacking. Although much of the article is basic common sense, and although I'm not a statistician, I had to seriously question the author's understanding of statistics at this point. The predetermined sample size (statistical power) is usually based on an assumption made about the effect size; i…