Live data from Hacker News

The Irreproducibility Crisis of Modern Science

nas.org

261–265 of 265 posts

Re: The Irreproducibility Crisis of Modern Science

#261
post #233
post #230

Earlier quoted context omitted.

Were I to run 100 independently designed experiments that all tested real effects, my choice of p-value does not determine the number that erroneously find no result. If the effect is small and I didn't gather enough data, a p-value of 0.05 could result in only a handful of experiments accurately reflecting reality. Let's say 30 make the cut. Were I to run another 100 independently designed experiments that all teste…

> Were I to run 100 independently designed experiments that all tested real effects, my choice of p-value does not determine the number that erroneously find no result. That's actually not true. The problem is that you cannot define what is a "real effect" without begging the question. Let me illustrate with an example: Suppose I do what appears to be a legitimate experiment to test a well-accepted law of nature. Unb…

> The problem is that you cannot define what is a "real effect" without begging the question

That's precisely my point. We don't know the sizes of these populations, so you really can't say anything quantitative about how the p-value affects the proportion of bad results in any given journal. My example was intentionally contrived.

You just have to be really careful when you talk about p-values because _so many people_ have this misunderstanding and it's actively harmful to getting the correct interpretation. That's why we keep going back and forth here — I'm not disagreeing that the situation is bad, I just want folks to recognize what the stats say and what they don't.

Re: The Irreproducibility Crisis of Modern Science

#262
post #261
post #233

Earlier quoted context omitted.

> Were I to run 100 independently designed experiments that all tested real effects, my choice of p-value does not determine the number that erroneously find no result. That's actually not true. The problem is that you cannot define what is a "real effect" without begging the question. Let me illustrate with an example: Suppose I do what appears to be a legitimate experiment to test a well-accepted law of nature. Unb…

> The problem is that you cannot define what is a "real effect" without begging the question That's precisely my point. We don't know the sizes of these populations, so you really can't say anything quantitative about how the p-value affects the proportion of bad results in any given journal. My example was intentionally contrived. You just have to be really careful when you talk about p-values because _so many peopl…

> We don't know the sizes of these populations

I think we can make some pretty reasonable guesses.

In any case, I certainly agree with you about the big picture: it's confusing, it's important to get it right, and a lot of people don't understand it. I may even be one of those people. But at the moment I believe that our disagreement is really over something else, namely, whether it's reasonable to have a prior on the null hypothesis.

Re: The Irreproducibility Crisis of Modern Science

#263

We know that weak classifiers can be bagged to produce a strong classifier al la Adaboost. Each study is a weak classifier and would have a 'reproducibility crisis' if retested on new data. However after lots of studies of similar phenomena, a strong classifier emerges. In the field, we call this converging lines of evidence.

Only if the classifiers have uncorrelated errors.

Re: The Irreproducibility Crisis of Modern Science

#264
post #259

Earlier quoted context omitted.

How can u get a wrong positive result if you only test correct H1s?

Because there is a very long causal chain between raw physical phenomena and your perceptions. Your equipment could be faulty. You could be suffering from hallucinations. You could be a brain in a vat. Example: I want to test if evolution is true. I hypothesize that if evolution is not true, then God will give me some sort of sign, e.g. I pray to God to make a coin that I flip come up heads if evolution is false. I f…

You are moving into metaphysics and including stuff that is not covered by Null Hypothesis Significance Testing. But on the original point "Likewise, in the long run, using a p threshold of 0.05 (which many journals do) will generate 5% false positive results" is just plain wrong. P(A|B) != P(A)

Re: The Irreproducibility Crisis of Modern Science

#265
post #259

Earlier quoted context omitted.

Because there is a very long causal chain between raw physical phenomena and your perceptions. Your equipment could be faulty. You could be suffering from hallucinations. You could be a brain in a vat. Example: I want to test if evolution is true. I hypothesize that if evolution is not true, then God will give me some sort of sign, e.g. I pray to God to make a coin that I flip come up heads if evolution is false. I f…

You are moving into metaphysics and including stuff that is not covered by Null Hypothesis Significance Testing. But on the original point "Likewise, in the long run, using a p threshold of 0.05 (which many journals do) will generate 5% false positive results" is just plain wrong. P(A|B) != P(A)

> P(A|B) != P(A)

It is if P(A) is 1, which is what you stipulated when you asked "How can u get a wrong positive result if you only test correct H1s?"

> "... using a p threshold of 0.05 ... will generate 5% false positive results" is just plain wrong

No, it's not wrong, it just makes a tacit assumption that most hypotheses that get experimentally tested are incorrect. Which is a reasonable assumption. If it were not true, making scientific progress would be a lot easier.

Post reply on HN