A glance at their new index shows that whatever they're measuring, it isn't useful.
Spend an hour with gemini 3.8 and tell me that model belongs in 2026. It feels like the model has Alzheimer's. It gets confused about whether what it reads is what it did. Just crazy bad.
I haven't tried muse spark 1.3. But it must have been a miracle since 1.2 to hit that rank.
Video game journalism vibes all over this.