For example, I found Kimi K3 to use more tokens than some other models, which caused it to cost twice as much purely because of the token volume. This experience lines up with ArtificialAnalysis’s benchmarks, but not these.
There’s a number of other comparisons here that don’t match up with my experience or other benchmarks. By many accounts, this is the outlier.
I could attribute the differences to harnesses used or something like reasoning levels, but none of those details are published.
While this seems interesting, I can’t take this seriously.
Correction: The harnesses are listed as a column, I missed that. My other concerns and questions still remain, it’s unclear why some of their results are the outlier that does not match my experience, ArtificalAnalysis’s benchmarks, or some of the experiences of others commenting.