Earlier quoted context omitted.
Every new proprietary model is "groundbreaking" and "look, it just solved task X that no other model could solve," only to be referred to as "that crappy previous-generation model" a month later. So yeah, I'm totally fine using Kimi-2.7, GLM-5.2 or Deepseek-v4. I think we've already hit the ceiling and most improvements now seem to be from harness improvements and slightly better RL to improve reasoning/tool calling.
There's also a lot of benchmark trickery going on, it's becoming harder to see how the latest models really improved. The top models also seem to have inconsistent performance depending on the time of day and how far we are from the next release.
https://marginlab.ai/trackers/claude-code-historical-perform...
There were at least a couple of these degradation trackers.