If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
- Sonnet 5 - 12.4%
- Luna - 17.3%
- Grok 4.6 - 20.3%
- Sol - 37.3%
- GLM 5.3 - 41.8%
- Opus 5 - 51.8%