Great article! Statistical fallacies are rampant in performance eval, even in academic settings. When designing statistical tests for performance, the keyword you want to use here is non-parametric. I.e., a U-test is a non-parametric analog to the t-test. It just looks at the rank statistics of results instead of their value, thus eliminating dependence on the underling distribution. Another issue that pops up is sam…
NIST has some nice simple descriptions and example experimental designs:
https://www.itl.nist.gov/div898/handbook/pmd/pmd.htm