More information on OpenAI's result (which seems better than DeepMind's) from the X thread: > our OpenAI reasoning system got a perfect score of 12/12 > For 11 of the 12 problems, the system’s first answer was correct. For the hardest problem, it succeeded on the 9th submission. Notably, the best human team achieved 11/12. > We had both GPT-5 and an experimental reasoning model generating solutions, and the experimen…
Ha so true. I was so tempted to copy and paste a problem into GPT5 and see what it would say
Hopefully that prompt was the same for all questions (I think that is what they did for the IMO submission, or maybe it was Google that did that, not sure).