If you're looking for a replacement you don't understand the problem. The problem isn't that P=.05 is an arbitrary measure of significance. The problem is that only publishing significant results is a bias against the null hypothesis . Let's say you're doing a study of flipping coins. The null hypothesis is that the coin is evenly weighted. If the null hypothesis is true, when you flip a coin once, it will come up he…
Is there a compendium somewhere of null hypotheses as well as observed p values (or whatever statistic) for experiments with both "significant" and "insignificant" results? Depending on the phenomenon and the hypothesis, an event could have a different probability of occurring and require a different threshold. It seems like we are wasting a lot of effort when any experiment is unpublished, when we could at the least…
You're absolutely right, but journals don't see it that way. Journals want to make money, and the articles which make the most money are the ones that prove an alternative hypothesis. This is one of the ways where for-profit publishing is harmful for science.
> This would require some standardization of hypotheses, so that a researcher could select a hypothesis from a list, conduct an experiment to test it, and report those findings to some aggregator. Others would also test the hypothesis and report their findings. Eventually you have some distribution of findings that allow you to examine the experimental methods of outliers as well as modal experiments. In this way every scientific result is the product of some meta-analysis, rather than allowing a single custom experiment to produce a result.
"Standardization of hypotheses" gets a bit tricky and I suspect the standardization process would be stifling. Part of the goal of replication is to improve on methodology, and part of that is finding ways in which your hypothesis wasn't clearly defined, or doesn't contribute to a larger theory. There needs to be some flexibility in which hypotheses scientists pursue.
A more organic way might be for journals to categorize articles as testing new hypotheses or attempting to reproduce old results, and strive for some ratio between the two (1:4 or somesuch).