Live data from Hacker News

Gemini with Deep Think achieves gold-medal standard at the IMO

deepmind.google

71–80 of 254 posts

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#71
post #61
post #6

Advanced Gemini, not Gemini Advanced. Thanks, Google. Maybe they should have named it MathBard.

They will have to rename gemini anyway, since it's doubtful they will ever be able to buy gemini.com. Gemini turboflex plus pro SE plaid interstellar 4.0

What makes you think google can't buy that domain?

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#72
post #25

> Btw as an aside, we didn’t announce on Friday because we respected the IMO Board's original request that all AI labs share their results only after the official results had been verified by independent experts & the students had rightly received the acclamation they deserved > We've now been given permission to share our results and are pleased to have been part of the inaugural cohort to have our model results off…

> Was OpenAI simply not coordinating with the IMO Board then? You are still surprised by sama@'s asinineness? You must be new here.

I am still surprised many people trust him. The board's (justified) decision to fire him was so awfully executed that it lead to him having even more slack

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#73
post #25

> Btw as an aside, we didn’t announce on Friday because we respected the IMO Board's original request that all AI labs share their results only after the official results had been verified by independent experts & the students had rightly received the acclamation they deserved > We've now been given permission to share our results and are pleased to have been part of the inaugural cohort to have our model results off…

This reminds me of when OpenAI made a splash (ages ago now) by beating the world's best Dota 2 teams using a RL model.

...Except they had to substantially bend the rules of the game (limiting the hero pool, completely changing/omitting certain mechanics) to pull this off. So they ended up beating some human Dota pros at a psuedo-Dota custom game, which was still impressive, but a very much watered-down result beneath the marketing hype.

It does seem like Money+Attention outweigh Science+Transparency at OpenAI, and this has always been the case.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#74
post #70

based on the score looks like they also couldn't answer question 6? has that been confirmed?

https://storage.googleapis.com/deepmind-media/gemini/IMO_202...

i saw that but it doesn't answer my question since it doesn't have associated marks? i'm not about to check their answer to a question i can't answer

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#75
post #25

> Btw as an aside, we didn’t announce on Friday because we respected the IMO Board's original request that all AI labs share their results only after the official results had been verified by independent experts & the students had rightly received the acclamation they deserved > We've now been given permission to share our results and are pleased to have been part of the inaugural cohort to have our model results off…

I think this is them not being confident enough before the event, so they don't wanna be shown a worse result than competitors. By being private they can obviously not publish anything if it didn't work out.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#77
post #8

Seems OpenAI knew this is forthcoming so they front ran the news? I was really high on Gemini 2.5 Pro after release but I kept going back to o3 for anything I cared about.

I regularly have the opposite experience: o3 is almost unusable, and Gemini 2.5 Pro is reliably great. Claude Opus 4 is a close second. o3 is so bad it makes me wonder if I'm being served a different model? My o3 responses are so truncated and simplified as to be useless. Maybe my problems aren't a good fit, but whatever it is: o3 output isn't useful.

Are you using a tool other than ChatGPT? If so, check the full prompt that's being sent. It can sometimes kneecap the model.

Tools having slightly unsuitable built in prompts/context sometimes lead to the models saying weird stuff out of the blue, instead of it actually being a 'baked in' behavior of the model itself. Seen this happen for both Gemini 2.5 Pro and o3.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#78

Do I understand it correctly that OpenAI self-proclaimed that they got their gold, without the official IMO judges grading their solutions?

Yes, OpenAI:

https://x.com/alexwei_/status/1946477754372985146

> 6/N In our evaluation, the model solved 5 of the 6 problems on the 2025 IMO. For each problem, three former IMO medalists independently graded the model’s submitted proof, with scores finalized after unanimous consensus. The model earned 35/42 points in total, enough for gold!

That means Google Deepmind is the first OFFICIAL IMO Gold.

https://x.com/demishassabis/status/1947337620226240803

> We've now been given permission to share our results and are pleased to have been part of the inaugural cohort to have our model results officially graded and certified by IMO coordinators and experts, receiving the first official gold-level performance grading for an AI system!

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#79
post #75
post #25

> Btw as an aside, we didn’t announce on Friday because we respected the IMO Board's original request that all AI labs share their results only after the official results had been verified by independent experts & the students had rightly received the acclamation they deserved > We've now been given permission to share our results and are pleased to have been part of the inaugural cohort to have our model results off…

I think this is them not being confident enough before the event, so they don't wanna be shown a worse result than competitors. By being private they can obviously not publish anything if it didn't work out.

As not-so-subtly hinted at by Terry Tao.

Its a great way to do PR but its a garbage way to to science.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#80
post #33

Still no information on the amount of compute needed; would be interested to see a breakdown from Google or OpenAI on what it took to achieve this feat. Something that was hotly debated in the thread with OpenAI's results: "We also provided Gemini with access to a curated corpus of high-quality solutions to mathematics problems, and added some general hints and tips on how to approach IMO problems to its instructions…

Ok but when reported by mass media, which never used SI units and instead uses units like libraries of Congress, or elephants, what kind of unit should media use to compare computational energy of ai vs children?

4.5 hours × 2 "days", 100 Wats including support system.

I'm not sure how to implement the "no calculator" rule :) but for this kind of problems it's not critical.

Total = 900Wh = 3.24MJ

Post reply on HN