Earlier quoted context omitted.
For the web app task I mentioned: * Kimi K3: 9532k input (9172k cached), 114k output - cost $5.5 * Qwen 3.8 Max: 18020k input (17836k cached), 114k output - cost $6.3 * Fable: ~14m input (all cached??), 196k output - cost $30 Correction on my earlier post, Kimi was through Pi, not Kimi Code. For Qwen I used Qwen Code and for Fable I used Claude Code. Not sure wtf is going on with the Fable stats (a lot tokens, virtua…
Would a fairer test not be to use the same harness for all three? I’d suspect the harness to massively affect token use and optimisation
Yes they do according to databricks -> https://www.databricks.com/blog/benchmarking-coding-agents-d...
That's why I find comparing models on benchmarks only gives the tendency. We should be comparing model x harness to have accurate metrics.