Earlier quoted context omitted.
These benchmark numbers are insane. The days when China was 6 months behind are over? How are they doing this with so much less resources than the US??? I have so much respect for the researchers there
To summarise the full results table further down the page (which doesn't render on the page for me!): Kimi K3 beats each model (out of 35 benchmarks, excluding missing): vs Fable 5 : 12/35 (34%) (ties: 1) vs GPT 5.6 Sol : 19/34 (56%) (ties: 1) vs Opus 4.8 : 30/35 (86%) vs GPT 5.5 : 30/34 (88%) (ties: 2) vs GLM-5.2 : 19/19 (100%) Beats Opus 4.8 and GPT 5.5 on all programming and agentic programming benchmarks except T…
Re: GLM-5.2: For a ~750b model, it holds up pretty good against models 3x its size (and ~10x the cost). Same goes for Tencent Hy3 and MiniMax M3, which almost match Opus 4.6 levels with ~295b params.