The problem with these benchmarks is that the Chinese models tend to be incredible on paper, and absolutely terrible in practice :/
I beg to differ. I replaced a $40/mo GitHub Copilot subscription where I used Opus 4.6 and GPT 5.5 with a $10/mo opencode Go plan where I use mostly DeepSeek V4 Flash and testing MiMo 2.5. I work on mid-sized projects currently (200k to 1kk lines of code).
Isn't that a million?