Live data from Hacker News

It’s not just p=0.048 vs. p=0.052

statmodeling.stat.columbia.edu

51–60 of 92 posts

Re: It’s not just p=0.048 vs. p=0.052

#51

Put a Number on It! did a piece a while ago going through some psychology pieces that were part of a replication effort. They found that half failed to replicate but that people in a betting market could often tell which ones were going to replicate or not. The author also did a blind test himself and was also able to guess which ones would replicate. He laid out several rules of thumb, most significantly to the arti…

That is a pretty interesting conclusion, do you know whether his work has been replicated to confirm the outcomes?

Re: It’s not just p=0.048 vs. p=0.052

#52
post #33

Related, about p-values: > Here's the problem in a nutshell: If you run 1000 experiments over the course of your career, and you get a significant effect (p > […] However, this is a statement about what happens when the null hypothesis is actually true. In real research, we don't know whether the null hypothesis is actually true. If we knew that, we wouldn't need any statistics! In real research, we have a p value, a…

That's not what the p-value means. It means that if you run 1000 of the experiments in a universe in which the hypothesis is false, around 50 of them will confirm the hypothesis anyway. If the hypothesis is true, then there are no false positives; all positives confirm the hypothesis. In a universe in which the hypothesis is true, there can only be false negatives.

"False positive" means that the effect or condition we're looking for is not true, but the experiment yields a true answer: the positive answer of the experiment is a falsehood. If the condition we're looking for is true, then there can't be a false positive. Even if the experiment yields a positive due to some flawed step, it's still a true positive.

Re: It’s not just p=0.048 vs. p=0.052

#53

Earlier quoted context omitted.

P-values are a sub-optimal but okay-ish of quantifying a Popperian hypothesis (a designed-to-be-refutable conjecture). The mathematics is not the problem, the problem is carving science (which in my view (and Quine's and others's) is pretty much defined by the unity of science) in testable morcels. None of the great achievements of science (Newton, Darwin, Mendeleyev, etc.) were obtained on the basis of Popperian dem…

Your great achievements in science exclude all real world applications (engineering, pharmaceuticals, etc), where the critical details of a hypothesis don't fit on a t-shirt

Note that I said "science", not "engineering". Engineering is driven by usefulness and profit.

Falsificationism isn't a stupid idea; it's even useful at a personal improvement level. But pharma or materials research use it because it tends to lead to good results, not because it's the very definition of what's worthwhile knowledge.

Re: It’s not just p=0.048 vs. p=0.052

#55
A 5% chance of the results being reproducible by fluke even if the hypothesis is false is obviously too high for anything important. Splitting hairs over .048 and .052 is ridiculous: it revolves around tiny differences in a gaping uncertainty. Neither value is anywhere in the neighborhood of where the benchmark should be.

Re: It’s not just p=0.048 vs. p=0.052

#56
It's worth noting that at their inception, null hypothesis testing and p-values were separate methodologies; there was even a bitter rivalry by their chief developers (Fisher and Pearson).

It wasn't until much later textbooks started to merge both. It may be worth to review Neyman and Pearson's attacks on Fisher in this matter.

Re: It’s not just p=0.048 vs. p=0.052

#57
post #33

Related, about p-values: > Here's the problem in a nutshell: If you run 1000 experiments over the course of your career, and you get a significant effect (p > […] However, this is a statement about what happens when the null hypothesis is actually true. In real research, we don't know whether the null hypothesis is actually true. If we knew that, we wouldn't need any statistics! In real research, we have a p value, a…

For about a year or so now Ive been wanting to make a game about science. It'd basically be a research and discovery simulator, and there would be a free-play mode. Some of the knobs would be # of required replication, required p-value and how good people are at generating hypothesis. I think it'd be eye opening.

There is a "people will mostly replicate/extend articles about X, and ignore articles about Y" (groupthink) effect that I imagine is also very relevant.

Re: It’s not just p=0.048 vs. p=0.052

#58
post #5

> To say it again: it is completely consistent with the null hypothesis to see p-values of 0.2 and 0.005 from two replications of the same damn experiment. I don't really follow this. Could someone clarify what is meant here? At what point would this author say something is not consistent with the null hypothesis?

The distribution of the p-value, given that the null is true, is uniform between 0 and 1.

IF the null is true, you're equally likely to get a p-value of 0.01 and 0.87.

Re: It’s not just p=0.048 vs. p=0.052

#59
post #5

> To say it again: it is completely consistent with the null hypothesis to see p-values of 0.2 and 0.005 from two replications of the same damn experiment. I don't really follow this. Could someone clarify what is meant here? At what point would this author say something is not consistent with the null hypothesis?

Its a meta observation of how we view p-values. He's saying that the same logic behind p-values and declaring something "statistically significant" when p And the argument will hold no matter what threshold you choose for rejecting the null hypothesis. You can choose to reject if p > X, and for any X, there will be values greater than X that, applying this meta-logic, are not statistically different from X.

> At what point would this author say something is not consistent with the null hypothesis?

Gelman's argument, I presume, is against the idea of significance testing as a whole. Declaring something "statistically significant" is in itself a very problematic thing, as it distills the entire phenomenon, the uncertainty surrounding the experiment, and the uncertainty surrounding the researcher's decisions to a single, binary conclusion.

Gelman is a Bayesian (perhaps the most famous modern Bayesian), and the Bayesian philosophy is to focus on producing a posterior distribution of the phenomenon being studied. I presume the alternative to significance and null hypothesis testing that he was suggest would be something closer to a model where people are reporting their priors/data/posteriors, and the discussion focuses around the implications and replication of those.

Re: It’s not just p=0.048 vs. p=0.052

#60
> Also, to get technical for a moment, the p-value is not the “probability of happening by chance.”

Is it not? According to Wikipedia, it's "[...] the probability that, when the null hypothesis is true, the statistical summary [...] would be equal to, or more extreme than, the actual observed results." This sounds pretty much like "probability of happening by chance".

Post reply on HN