Earlier quoted context omitted.
Why? These models just leapfrog each other as time advances. One month Gemini is on top, then ChatGPT, then Anthropic. Not sure why everyone gets FOMO whenever a new version gets released.
Considering GPT 5 was only recently released, it's very unlikely GPT will achieve these scores in just a couple of months. If they had something this good in the oven, they'd probably left the GPT 5 name to it. Or maybe Google just benchmaxxed and this doesn't translate at all in real world performance.
TBD if that performance generalizes to other real world tasks.