Live data from Hacker News

Gemini with Deep Think achieves gold-medal standard at the IMO

deepmind.google

131–140 of 254 posts

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#131
post #97

Earlier quoted context omitted.

I found the proofs you were referring to: Google https://storage.googleapis.com/deepmind-media/gemini/IMO_202... OpenAI https://github.com/aw31/openai-imo-2025-proofs/

Gemini is clearer but MY GOD is it verbose. e.g. look at problem 1, section 2. Analysis of the Core Problem - there's nothing at all deep here, but it seems the model wants to spell out every single tiny logical step. I wonder if this is a stylistic choice or something that actually helps the model get to the end.

Section 2 is a case by case analysis. Those are never pretty but perfectly normal given the problem.

With OpenAI that part takes up about 2/3 if the proof even with its fragmented prose. I don't think it does much better.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#132

From Terence Tao, via mastodon [0]: > It is tempting to view the capability of current AI technology as a singular quantity: either a given task X is within the ability of current tools, or it is not. However, there is in fact a very wide spread in capability (several orders of magnitude) depending on what resources and assistance gives the tool, and how one reports their results. > One can illustrate this with a hum…

Some of the critique is valid but some of it sounds like, "but the rules of the contest are that participants must use less than x joules of energy obtained from cellular respiration and have a singular consciousness" I don't think anybody thinks AI was competing fair and within the rules that apply to humans. But if the humans were competing on the terms that AI solved those problems on, near-unlimited access to ene…

I don't think that characterization is fair at all. It's certainly true that you, me, and most humans can't solve these problems with any amount of time or energy. But the problems are specifically written to be at the limit of what the actual high school students who participate can solve in four hours. Letting the actual students taking the test have four days instead of four hours would make a massive difference in their ability to solve them.

Said differently, the students, difficulty of the problems, and time limit are specifically coordinated together, so the amount of joules of energy used to produce a solution is not arbitrary. In the grand scheme of how the tech will improve over time, it seems likely that doesn't matter and the computers will win by any metric soon enough, but Tao is completely correct to point out that you haven't accurately told us what the machines can do today, in July 2025, without telling us ahead of time exactly what rules you are modifying.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#133

I think we are having a Deep Blue vs. Kasparov moment in Competitive Math right now. This is a large progress from just a few years ago and yet I think we still are really far away from even a semi-respectable AI mathematician. What an exciting time to be alive!

More like Deep Blue vs Child prodigy. In the IMO the contestants are high school students not the greatest mathematicians in the world.

Of course contest math is not research math but the active IMO kids are pretty much the best in the world at math contests.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#134
Woah they used parallel reasoning. An idea I opensourced about a month before GDMs first paper on it. Very cool. https://x.com/GoogleDeepMind/status/1947333836594946337 So you might be able to achieve similar performance at home today using llm-consortium https://github.com/irthomasthomas/llm-consortium

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#135
post #106

Earlier quoted context omitted.

Ah thanks that does answer it. You should start a blog

Simon's a prolific writer on AI http://simonw.substack.com https://simonwillison.net

cool! glad he took my advice

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#136
> all within the 4.5-hour competition time limit

Both OpenAI and Google pointed this out, but does that matter a lot? They could have spun up a million parallel reasoning processes to search for a proof that checks out - though of course some large amount of computation would have to be reserved for some kind of evaluator model to rank the proofs and decide which one to submit. Perhaps it was hundreds of years of GPU time.

Though of course it remains remarkable that this kind of process finds solutions at all and is even parallelizable to this degree, perhaps that is what they meant. And I also don't want to diminish the significance of the result, since in the end it doesn't matter if we get AGI with overwhelming compute or not. The human brain doesn't scale as nicely, even if it's more energy efficient.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#137
post #87

Earlier quoted context omitted.

Sounds like it did not: > This year, our advanced Gemini model operated end-to-end in natural language, producing rigorous mathematical proofs directly from the official problem descriptions – all within the 4.5-hour competition time limit

I interpreted that bit as meaning they did not manually alter the problem statement before feeding it to the model - they gave it the exact problem text issued by IMO. It is not clear to me from that paragraph if the model was allowed to call tools on its own or not.

As a side question, do you think using tools like Lean will become a staple of these "deep reasoning" LLM flavors?

It seems that LLMs excel (relative to other paradigms) in the kind of "loose" creative thinking humans do, but are also prone to the same kinds of mistakes humans make (hallucinations, etc). Just as Lean and other formal systems can help humans find subtle errors in their own thinking, they could do the same for LLMs.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#139
Most critical piece of information I couldn’t find is - how many shot was this?

Could it understand the solution is correct by itself (one-shot)? Or did it have just great math intuition and knowledge? How the solutions were validated if it was 10-100 shot?

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#140
post #88

Earlier quoted context omitted.

What a great metaphor for AI. Taking an event that is a celebration of high school kids' knowledge and abilities and turning it into a marketing stunt for their frankenstein monster that they are building to make all the kids' hard work worth nothing.

Not only, by not officially entering they had no obligation to announce their result so if they didn't achieve a gold medal score they presumably wouldn't have made any announcement and no-one would have been the wiser. This cowardly bullshit followed by the grandstanding on Twitter is high-school bully behaviour.

If they failed and remained quiet, then everyone would know that the other companies performed well and they didn't even qualify.
Post reply on HN