I think I have a minority opinion here, but I’m a bit disappointed they seem to be moving away from formal techniques. I think if you ever want to truly “automate” math or do it at machine scale, e.g. creating proofs that would amount to thousands of pages of writing, there is simply no way forward but to formalize. Otherwise, one cannot get past the bottleneck of needing a human reviewer to understand and validate the proof.
Gemini with Deep Think achieves gold-medal standard at the IMO
91–100 of 254 posts
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#92Surprising since Reid Barton was working on a lean system.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#93Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#94Earlier quoted context omitted.
Yes, OpenAI: https://x.com/alexwei_/status/1946477754372985146 > 6/N In our evaluation, the model solved 5 of the 6 problems on the 2025 IMO. For each problem, three former IMO medalists independently graded the model’s submitted proof, with scores finalized after unanimous consensus. The model earned 35/42 points in total, enough for gold! That means Google Deepmind is the first OFFICIAL IMO Gold. https://x.com/demi…
Do you know if OpenAI used the same grading criteria as official judges?
But this can be verified because the results are public:
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#95Seems OpenAI knew this is forthcoming so they front ran the news? I was really high on Gemini 2.5 Pro after release but I kept going back to o3 for anything I cared about.
>I was really high on Gemini 2.5 Pro after release but I kept going back to o3 for anything I cared about Same here. I was impressed by their benchmarks and topping most leaderboards, but in day to day use they still feel so far behind.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#96Earlier quoted context omitted.
I think this is them not being confident enough before the event, so they don't wanna be shown a worse result than competitors. By being private they can obviously not publish anything if it didn't work out.
As not-so-subtly hinted at by Terry Tao. Its a great way to do PR but its a garbage way to to science.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#97Comparing the answers between Openai and Gemini the writing style of Gemini is a lot clearer. It could be presented a bit better but it's easy enough to follow the proof. This also makes it a lot shorter than the answer given by OpenAI and it uses proper prose.
Google https://storage.googleapis.com/deepmind-media/gemini/IMO_202...
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#98Earlier quoted context omitted.
As not-so-subtly hinted at by Terry Tao. Its a great way to do PR but its a garbage way to to science.
True, but openai definitely isn't trying to do public research on science, they are all about money now.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#99Earlier quoted context omitted.
i saw that but it doesn't answer my question since it doesn't have associated marks? i'm not about to check their answer to a question i can't answer
That PDF lists solutions for problems 1 through 5 but does not mention problem 6 at all.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#100Earlier quoted context omitted.
I regularly have the opposite experience: o3 is almost unusable, and Gemini 2.5 Pro is reliably great. Claude Opus 4 is a close second. o3 is so bad it makes me wonder if I'm being served a different model? My o3 responses are so truncated and simplified as to be useless. Maybe my problems aren't a good fit, but whatever it is: o3 output isn't useful.
I have this distinctive feeling that o3 tries to trick me intentionally when it can't solve a problem by cleverly hiding its mistakes. But I could be imagining it