Live data from Hacker News

Gemini with Deep Think achieves gold-medal standard at the IMO

deepmind.google

61–70 of 254 posts

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#61
post #6

Advanced Gemini, not Gemini Advanced. Thanks, Google. Maybe they should have named it MathBard.

They will have to rename gemini anyway, since it's doubtful they will ever be able to buy gemini.com. Gemini turboflex plus pro SE plaid interstellar 4.0

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#62
post #24

Related news: - OpenAI claims gold-medal performance at IMO 2025 https://news.ycombinator.com/item?id=44613840 - "According to a friend, the IMO asked AI companies not to steal the spotlight from kids and to wait a week after the closing ceremony to announce results. OpenAI announced the results BEFORE the closing ceremony. According to a Coordinator on Problem 6, the one problem OpenAI couldn't solve, "the general s…

What a great metaphor for AI. Taking an event that is a celebration of high school kids' knowledge and abilities and turning it into a marketing stunt for their frankenstein monster that they are building to make all the kids' hard work worth nothing.

Google did the correct and respectful thing.

It appears that OpenAI didn't officially enter (whereas Google did), that they knew Google was going to gold medal, and that they released their news ahead of time (disrespecting the kids and organizers) so they could scoop Google.

Really scummy on OpenAI's part.

The IMO closing ceremony was July 19. OpenAI announced on the same day.

IMO requested the tech companies to wait until the following week.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#63

> AlphaGeometry and AlphaProof required experts to first translate problems from natural language into domain-specific languages, such as Lean, and vice-versa for the proofs. It also took two to three days of computation. This year, our advanced Gemini model operated end-to-end in natural language, producing rigorous mathematical proofs directly from the official problem descriptions So, the problem wasn't translated…

Sounds like it did not:

> This year, our advanced Gemini model operated end-to-end in natural language, producing rigorous mathematical proofs directly from the official problem descriptions – all within the 4.5-hour competition time limit

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#64
post #15

Earlier quoted context omitted.

Does the IMO reuse problems? My understanding is that new problems are submitted each year and 6 are selected for each competition. The submitted problems are then published after the IMO has concluded. How would the training data contain unpublished, newly submitted problems? Obviously the training data contained similar problems, because that's what every IMO participant already studies. It seems unlikely that they…

IMO doesn't reuse problems, but Terence Tao has a Mastodon post where he explains that the first five (of six) problems are generally ones where existing techniques can be leveraged to get to the answer. The sixth problem requires considerable originality. Notably, both Gemini and OpenAI's model didn't get the sixth problem. Still quite an achievement though.

Do you have another source for that? I checked his Mastodon feed and don't see any mention about the source of the questions from the IMO.

https://mathstodon.xyz/@tao

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#65
post #6

Advanced Gemini, not Gemini Advanced. Thanks, Google. Maybe they should have named it MathBard.

Google and Microsoft continuing to prove that the hardest problem in programming is naming things.

There are two different versions of that hard problem: the computer science version, and the marketing version. The marketing one has more nebulous acceptance criteria though.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#66
post #8

Seems OpenAI knew this is forthcoming so they front ran the news? I was really high on Gemini 2.5 Pro after release but I kept going back to o3 for anything I cared about.

I regularly have the opposite experience: o3 is almost unusable, and Gemini 2.5 Pro is reliably great. Claude Opus 4 is a close second. o3 is so bad it makes me wonder if I'm being served a different model? My o3 responses are so truncated and simplified as to be useless. Maybe my problems aren't a good fit, but whatever it is: o3 output isn't useful.

I have this distinctive feeling that o3 tries to trick me intentionally when it can't solve a problem by cleverly hiding its mistakes. But I could be imagining it

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#67

Still no information on the amount of compute needed; would be interested to see a breakdown from Google or OpenAI on what it took to achieve this feat. Something that was hotly debated in the thread with OpenAI's results: "We also provided Gemini with access to a curated corpus of high-quality solutions to mathematics problems, and added some general hints and tips on how to approach IMO problems to its instructions…

Some unofficial comparison with costs of public models (performing worse): https://matharena.ai/imo/

So the real cost is something much more.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#69
post #33

Still no information on the amount of compute needed; would be interested to see a breakdown from Google or OpenAI on what it took to achieve this feat. Something that was hotly debated in the thread with OpenAI's results: "We also provided Gemini with access to a curated corpus of high-quality solutions to mathematics problems, and added some general hints and tips on how to approach IMO problems to its instructions…

Ok but when reported by mass media, which never used SI units and instead uses units like libraries of Congress, or elephants, what kind of unit should media use to compare computational energy of ai vs children?

Dollars of compute at market rate is what I'd like to see, to check whether calling this tool would cost $100 or $100,000
Post reply on HN