We know that weak classifiers can be bagged to produce a strong classifier al la Adaboost. Each study is a weak classifier and would have a 'reproducibility crisis' if retested on new data. However after lots of studies of similar phenomena, a strong classifier emerges. In the field, we call this converging lines of evidence.
I'm not sure I really agree with this, because in science studies are biased towards aligning with the results of previous studies. This can be due to many factors: authors' expectations, journals like to publish positive over negative results, peer reviewers will look more critically (in an analytical sense) at work that disagrees with accepted research, etc.... It can be very hard for many reasons to stand up and s…
To me it brings up an interesting point. If we view experiments as classifiers, then how would a machine learning expert set science policy and practice?
Also viewed this way, it makes me wonder about p-hacking, which increases sensitivity while increasing false positives. Since negative results are not generally reported, I wonder whether p-hacking diminishes replicability at the study level for efficiency at the aggregate level. This is an empirical question, as ethically it is of course a dubious practice under current understanding.