Advanced Gemini, not Gemini Advanced. Thanks, Google. Maybe they should have named it MathBard.
They will have to rename gemini anyway, since it's doubtful they will ever be able to buy gemini.com. Gemini turboflex plus pro SE plaid interstellar 4.0
Gemini with Deep Think achieves gold-medal standard at the IMO
71–80 of 254 posts
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#72> Btw as an aside, we didn’t announce on Friday because we respected the IMO Board's original request that all AI labs share their results only after the official results had been verified by independent experts & the students had rightly received the acclamation they deserved > We've now been given permission to share our results and are pleased to have been part of the inaugural cohort to have our model results off…
> Was OpenAI simply not coordinating with the IMO Board then? You are still surprised by sama@'s asinineness? You must be new here.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#73> Btw as an aside, we didn’t announce on Friday because we respected the IMO Board's original request that all AI labs share their results only after the official results had been verified by independent experts & the students had rightly received the acclamation they deserved > We've now been given permission to share our results and are pleased to have been part of the inaugural cohort to have our model results off…
...Except they had to substantially bend the rules of the game (limiting the hero pool, completely changing/omitting certain mechanics) to pull this off. So they ended up beating some human Dota pros at a psuedo-Dota custom game, which was still impressive, but a very much watered-down result beneath the marketing hype.
It does seem like Money+Attention outweigh Science+Transparency at OpenAI, and this has always been the case.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#74based on the score looks like they also couldn't answer question 6? has that been confirmed?
https://storage.googleapis.com/deepmind-media/gemini/IMO_202...
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#75> Btw as an aside, we didn’t announce on Friday because we respected the IMO Board's original request that all AI labs share their results only after the official results had been verified by independent experts & the students had rightly received the acclamation they deserved > We've now been given permission to share our results and are pleased to have been part of the inaugural cohort to have our model results off…
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#76I will be surprised when a model with only the knowledge of a college student can solve these problems.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#77Seems OpenAI knew this is forthcoming so they front ran the news? I was really high on Gemini 2.5 Pro after release but I kept going back to o3 for anything I cared about.
I regularly have the opposite experience: o3 is almost unusable, and Gemini 2.5 Pro is reliably great. Claude Opus 4 is a close second. o3 is so bad it makes me wonder if I'm being served a different model? My o3 responses are so truncated and simplified as to be useless. Maybe my problems aren't a good fit, but whatever it is: o3 output isn't useful.
Tools having slightly unsuitable built in prompts/context sometimes lead to the models saying weird stuff out of the blue, instead of it actually being a 'baked in' behavior of the model itself. Seen this happen for both Gemini 2.5 Pro and o3.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#78Do I understand it correctly that OpenAI self-proclaimed that they got their gold, without the official IMO judges grading their solutions?
https://x.com/alexwei_/status/1946477754372985146
> 6/N In our evaluation, the model solved 5 of the 6 problems on the 2025 IMO. For each problem, three former IMO medalists independently graded the model’s submitted proof, with scores finalized after unanimous consensus. The model earned 35/42 points in total, enough for gold!
That means Google Deepmind is the first OFFICIAL IMO Gold.
https://x.com/demishassabis/status/1947337620226240803
> We've now been given permission to share our results and are pleased to have been part of the inaugural cohort to have our model results officially graded and certified by IMO coordinators and experts, receiving the first official gold-level performance grading for an AI system!
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#79> Btw as an aside, we didn’t announce on Friday because we respected the IMO Board's original request that all AI labs share their results only after the official results had been verified by independent experts & the students had rightly received the acclamation they deserved > We've now been given permission to share our results and are pleased to have been part of the inaugural cohort to have our model results off…
I think this is them not being confident enough before the event, so they don't wanna be shown a worse result than competitors. By being private they can obviously not publish anything if it didn't work out.
Its a great way to do PR but its a garbage way to to science.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#80Still no information on the amount of compute needed; would be interested to see a breakdown from Google or OpenAI on what it took to achieve this feat. Something that was hotly debated in the thread with OpenAI's results: "We also provided Gemini with access to a curated corpus of high-quality solutions to mathematics problems, and added some general hints and tips on how to approach IMO problems to its instructions…
Ok but when reported by mass media, which never used SI units and instead uses units like libraries of Congress, or elephants, what kind of unit should media use to compare computational energy of ai vs children?
I'm not sure how to implement the "no calculator" rule :) but for this kind of problems it's not critical.
Total = 900Wh = 3.24MJ