Advanced Gemini, not Gemini Advanced. Thanks, Google. Maybe they should have named it MathBard.
Gemini with Deep Think achieves gold-medal standard at the IMO
61–70 of 254 posts
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#62Related news: - OpenAI claims gold-medal performance at IMO 2025 https://news.ycombinator.com/item?id=44613840 - "According to a friend, the IMO asked AI companies not to steal the spotlight from kids and to wait a week after the closing ceremony to announce results. OpenAI announced the results BEFORE the closing ceremony. According to a Coordinator on Problem 6, the one problem OpenAI couldn't solve, "the general s…
What a great metaphor for AI. Taking an event that is a celebration of high school kids' knowledge and abilities and turning it into a marketing stunt for their frankenstein monster that they are building to make all the kids' hard work worth nothing.
It appears that OpenAI didn't officially enter (whereas Google did), that they knew Google was going to gold medal, and that they released their news ahead of time (disrespecting the kids and organizers) so they could scoop Google.
Really scummy on OpenAI's part.
The IMO closing ceremony was July 19. OpenAI announced on the same day.
IMO requested the tech companies to wait until the following week.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#63> AlphaGeometry and AlphaProof required experts to first translate problems from natural language into domain-specific languages, such as Lean, and vice-versa for the proofs. It also took two to three days of computation. This year, our advanced Gemini model operated end-to-end in natural language, producing rigorous mathematical proofs directly from the official problem descriptions So, the problem wasn't translated…
> This year, our advanced Gemini model operated end-to-end in natural language, producing rigorous mathematical proofs directly from the official problem descriptions – all within the 4.5-hour competition time limit
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#64Earlier quoted context omitted.
Does the IMO reuse problems? My understanding is that new problems are submitted each year and 6 are selected for each competition. The submitted problems are then published after the IMO has concluded. How would the training data contain unpublished, newly submitted problems? Obviously the training data contained similar problems, because that's what every IMO participant already studies. It seems unlikely that they…
IMO doesn't reuse problems, but Terence Tao has a Mastodon post where he explains that the first five (of six) problems are generally ones where existing techniques can be leveraged to get to the answer. The sixth problem requires considerable originality. Notably, both Gemini and OpenAI's model didn't get the sixth problem. Still quite an achievement though.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#65Advanced Gemini, not Gemini Advanced. Thanks, Google. Maybe they should have named it MathBard.
Google and Microsoft continuing to prove that the hardest problem in programming is naming things.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#66Seems OpenAI knew this is forthcoming so they front ran the news? I was really high on Gemini 2.5 Pro after release but I kept going back to o3 for anything I cared about.
I regularly have the opposite experience: o3 is almost unusable, and Gemini 2.5 Pro is reliably great. Claude Opus 4 is a close second. o3 is so bad it makes me wonder if I'm being served a different model? My o3 responses are so truncated and simplified as to be useless. Maybe my problems aren't a good fit, but whatever it is: o3 output isn't useful.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#67Still no information on the amount of compute needed; would be interested to see a breakdown from Google or OpenAI on what it took to achieve this feat. Something that was hotly debated in the thread with OpenAI's results: "We also provided Gemini with access to a curated corpus of high-quality solutions to mathematics problems, and added some general hints and tips on how to approach IMO problems to its instructions…
So the real cost is something much more.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#68Advanced Gemini, not Gemini Advanced. Thanks, Google. Maybe they should have named it MathBard.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#69Still no information on the amount of compute needed; would be interested to see a breakdown from Google or OpenAI on what it took to achieve this feat. Something that was hotly debated in the thread with OpenAI's results: "We also provided Gemini with access to a curated corpus of high-quality solutions to mathematics problems, and added some general hints and tips on how to approach IMO problems to its instructions…
Ok but when reported by mass media, which never used SI units and instead uses units like libraries of Congress, or elephants, what kind of unit should media use to compare computational energy of ai vs children?
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#70based on the score looks like they also couldn't answer question 6? has that been confirmed?