Earlier quoted context omitted.
> I just don't believe non-deterministic tools can actually be benchmarked. It's all hoopla to me. We benchmark non-deterministic things all the time and it's frankly not even that unusual or hard. You yourself indicate that one model outperforms another one in your experience on various facets, and that is itself a benchmark. The more relevant question is probably how well does a given benchmark translate to improve…
My anecdotal experience isn't a benchmark. Just because I feel like something is better or different doesn't mean it actually is.
Of course, but it is a data point, and multiple such data points can be aggregated. This is true even if all you can do is compare two things.
The shape of that data will reveal something more about the thing you want to measure than the null hypothesis you'd otherwise have.