If you're looking for a replacement you don't understand the problem. The problem isn't that P=.05 is an arbitrary measure of significance. The problem is that only publishing significant results is a bias against the null hypothesis . Let's say you're doing a study of flipping coins. The null hypothesis is that the coin is evenly weighted. If the null hypothesis is true, when you flip a coin once, it will come up he…
It seems like we are wasting a lot of effort when any experiment is unpublished, when we could at the least be associating a data point with a particular hypothesis.
This would require some standardization of hypotheses, so that a researcher could select a hypothesis from a list, conduct an experiment to test it, and report those findings to some aggregator. Others would also test the hypothesis and report their findings. Eventually you have some distribution of findings that allow you to examine the experimental methods of outliers as well as modal experiments. In this way every scientific result is the product of some meta-analysis, rather than allowing a single custom experiment to produce a result.
The current approach also ties the experimental method to the result, requiring potentially fallible oversight of the experimental design. The hypothesis-testing aggregator reduces this reliance, which is perhaps more honest about our ability to design good experiments, would allow for a greater diversity of approaches, and could also be more efficient when evaluating results.
The most interesting aspect of this is if you could somehow create relationships between standardized hypotheses found on the hypothesis menu such that you could infer other hypotheses in a more systematic way.