Gemini with Deep Think achieves gold-medal standard at the IMO
31–40 of 254 posts
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#32Do I understand it correctly that OpenAI self-proclaimed that they got their gold, without the official IMO judges grading their solutions?
Goog had an official colab with IMO, and we can be sure they got those results under the imposed constraints (last year they allocated ~48h for silver IIRC) and an official grading by the IMO graders.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#33Still no information on the amount of compute needed; would be interested to see a breakdown from Google or OpenAI on what it took to achieve this feat. Something that was hotly debated in the thread with OpenAI's results: "We also provided Gemini with access to a curated corpus of high-quality solutions to mathematics problems, and added some general hints and tips on how to approach IMO problems to its instructions…
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#34> Btw as an aside, we didn’t announce on Friday because we respected the IMO Board's original request that all AI labs share their results only after the official results had been verified by independent experts & the students had rightly received the acclamation they deserved > We've now been given permission to share our results and are pleased to have been part of the inaugural cohort to have our model results off…
You are still surprised by sama@'s asinineness? You must be new here.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#35I think we are having a Deep Blue vs. Kasparov moment in Competitive Math right now. This is a large progress from just a few years ago and yet I think we still are really far away from even a semi-respectable AI mathematician. What an exciting time to be alive!
Your comparison with chess engines is pretty spot-on, that's how the best of the best chess players do prep nowadays. Gone are the multi person expert teams that analysed positions and offered advice. They now have analysts that use supercomputers to search through bajillions of positions and extract the best ideas, and distill them to their players.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#36> Btw as an aside, we didn’t announce on Friday because we respected the IMO Board's original request that all AI labs share their results only after the official results had been verified by independent experts & the students had rightly received the acclamation they deserved > We've now been given permission to share our results and are pleased to have been part of the inaugural cohort to have our model results off…
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#37Do I understand it correctly that OpenAI self-proclaimed that they got their gold, without the official IMO judges grading their solutions?
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#38Related news: - OpenAI claims gold-medal performance at IMO 2025 https://news.ycombinator.com/item?id=44613840 - "According to a friend, the IMO asked AI companies not to steal the spotlight from kids and to wait a week after the closing ceremony to announce results. OpenAI announced the results BEFORE the closing ceremony. According to a Coordinator on Problem 6, the one problem OpenAI couldn't solve, "the general s…
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#39Do I understand it correctly that OpenAI self-proclaimed that they got their gold, without the official IMO judges grading their solutions?
Childish. And, of course they must have known there was an official LLM cohort taking the real test, and they probably even knew that Gemini got a gold medal, and may have even known that Google planned a press release for today.
> We were trying to get a big client for weeks, and they said no and went with a competitor. The competitor already had a terms sheet from the company were we trying to sign up. It was real serious.
> We were devastated, but we decided to fly down and sit in their lobby until they would meet with us. So they finally let us talk to them after most of the day.
> We then had a few more meetings, and the company wanted to come visit our offices so they could make sure we were a 'real' company. At that time, we were only 5 guys. So we hired a bunch of our college friends to 'work' for us for the day so we could look larger than we actually were. It worked, and we got the contract.
> I think the reason why PG respects Sam so much is he is charismatic, resourceful, and just overall seems like a genuine person.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#40Earlier quoted context omitted.
>I was really high on Gemini 2.5 Pro after release but I kept going back to o3 for anything I cared about Same here. I was impressed by their benchmarks and topping most leaderboards, but in day to day use they still feel so far behind.
I think that's most likely just your view, and not really based on evidence.