Live data from Hacker News

Gemini with Deep Think achieves gold-medal standard at the IMO

deepmind.google

111–120 of 254 posts

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#112
How much of a big deal is this stuff? I was blessed with dyscalculia so I can hardly add two numbers together, don't pay much attention to the mathematics word, but my reading indicates this is extremely difficult/humans cannot do this?

What comes next for this particular exercise? Thank you.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#113

Earlier quoted context omitted.

This is a fair reply, but TBH I don't think it's going to change much. The upper echelon of the human society has decided to move AI forward rapidly regardless of any consequences. The rest of us can only hold and pray.

You are watching american money hard at work, my friend. It's either glorious or reckless, hard to tell for now.

Could be both, but one for different group of people.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#115
Useful and interesting but likely still dangerous in production without connecting to formal verification tools.

I know o3 is far from state of the art these days but it's great at finding relevant literature and suggesting inequalities to consider but in actual proofs it can produce convincing looking statements that are false if you follow the details, or even just the algebra, carefully. Subtle errors like these might become harder to detect as the models get better.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#117
post #75
post #25

> Btw as an aside, we didn’t announce on Friday because we respected the IMO Board's original request that all AI labs share their results only after the official results had been verified by independent experts & the students had rightly received the acclamation they deserved > We've now been given permission to share our results and are pleased to have been part of the inaugural cohort to have our model results off…

I think this is them not being confident enough before the event, so they don't wanna be shown a worse result than competitors. By being private they can obviously not publish anything if it didn't work out.

They shot themselves in the foot by not showing the confidence that Google did.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#118

Solved 5 problems out of 6, scoring 35 out of 42. For comparison OpenAI scored 35/42 too, days back.

I wouldn't read too much into the timelines, as it seems that OpenAI simply broke an embargo that the other players were up to that point respecting: https://arstechnica.com/ai/2025/07/openai-jumps-gun-on-inter...

Very in character for them!

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#119

Earlier quoted context omitted.

>it seems that the answer to whether or not a general model could perform such a feat is that the models were trained specifically on IMO problems, which is what a number of folks expected. Not sure thats exactly what that means. Its already likely the case that these models contained IMO problems and solutions from pretraining. It's possible this means they were present in the system prompt or something similar.

Does the IMO reuse problems? My understanding is that new problems are submitted each year and 6 are selected for each competition. The submitted problems are then published after the IMO has concluded. How would the training data contain unpublished, newly submitted problems? Obviously the training data contained similar problems, because that's what every IMO participant already studies. It seems unlikely that they…

>How would the training data contain unpublished, newly submitted problems?

I don't think I or op suggested it did.

Post reply on HN