Live data from Hacker News

Gemini with Deep Think achieves gold-medal standard at the IMO

deepmind.google

221–230 of 254 posts

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#221

Earlier quoted context omitted.

If a Language Model is capable of producing rigorous natural language proofs then getting it to produce Lean (or whatever) proofs would not be a big deal. This is a wildly uninformed take. Even today there are plenty of basic statements which LLM’s can produce English language proofs of that have not been formalized.

Most mathematicians aren't much interested in translating statements for the fun of it, so whether a lot of basic statements are un-formalized doesn't mean much. And the point was never that formalization was easy. I said that if language models become capable enough of not needing a crutch, then adding one afterwards isn't a big deal. What exactly do you think Alphaproof is? Much worse LLMs were already doing what y…

AlphaProof works natively in Lean. Literally the one part it couldn’t do was translate the natural language statements into the formalized language; that was done manually by humans.

I’m saying that the view of formalized languages as a crutch is backwards; they are in fact a challenging constraint.

The most extreme version of my position is that there is no such thing as a rigorous natural language proof, and that humans will view the relative informality of 20th and early 21st century mathematics much like we currently view the informality of pre-18th century mathematics.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#222

Earlier quoted context omitted.

Most mathematicians aren't much interested in translating statements for the fun of it, so whether a lot of basic statements are un-formalized doesn't mean much. And the point was never that formalization was easy. I said that if language models become capable enough of not needing a crutch, then adding one afterwards isn't a big deal. What exactly do you think Alphaproof is? Much worse LLMs were already doing what y…

AlphaProof works natively in Lean. Literally the one part it couldn’t do was translate the natural language statements into the formalized language; that was done manually by humans. I’m saying that the view of formalized languages as a crutch is backwards; they are in fact a challenging constraint. The most extreme version of my position is that there is no such thing as a rigorous natural language proof, and that h…

I did not say formal languages are a crutch. I said they served as a crutch for the Alphaproof system, because they literally did.

It was there so the generator wouldn't go off the rails hallucinating and candidates could be verified. And the solving still took days. Take that away and the generator doesn't get anywhere near silver with all the time in the world.

That is the definition of a crutch. The generator was smart enough to produce promising candidates with enough brute force trial and error, but that's it. Without Lean, it wouldn't have worked.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#223

Earlier quoted context omitted.

this comment reminds me of that Feynman quote about others thinking scientific knowledge removes the beauty from a flower. of course Feynman disagreed

Feynman isn't a god and he didn't conceive of AI.

Von Neumann pretty much was a god, and even as early as 1945, he was explicitly assuming that neural models were how things in the computing business would eventually shake out.

The rest of us are just a little slow to catch up, is all.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#224
post #212
post #207

Earlier quoted context omitted.

> My understanding is that they did, but don't any more; it's no longer true that humans understand enough things about chess better than computers for the human/computer collaboration to contribute anything over just using the computer. This is not true, at least not in very long time formats like correspondence chess: https://en.chessbase.com/post/correspondence-chess-and-corre... There's also many well known cases…

Your Nakamura example is from 2008. That's 17 years ago. The machines have improved a lot since then, hardware and software both. I've seen Nakamura beat strong-but-still-limited bots playing "anti-computer chess" but I am fairly sure he would be eaten alive if he tried it against present-day Stockfish or Leela on good hardware. Maybe you're right about correspondence chess. That interview is from 2018 and the machin…

I'm aware the Nakamura example is old, but the core issue (the horizon effect) is still there in any alpha/beta pruning engine, including the newest SF. But I will certainly grant you that it has become much harder to execute since then. MCTS engines (like Lc0) are far more immune to the horizon effect, but can instead suffer from missing very shallow tactics, especially in very fast time controls, as Andrew Tang showed in his match against an early version of Lc0 7 years ago: https://lichess.org/@/lichess/blog/gm-andrew-tang-defends-hu...

> Maybe you're right about correspondence chess. That interview is from 2018 and the machines have got distinctly stronger in that time, but 7 years isn't so long and it could be that human input still has some value for CC.

I've played on ICCF and I can say that the draw situation is much worse now, exactly because engines are so much stronger now, so it's much harder to find good opening novelties against a competent opponent. Engines were infamous for misjudging many closed openings like the KID, but even in the last few years that has really tightened up.

Still, the human element is useful in trying to steer clear of drawish openings and middlegame lines.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#225
post #175
post #136

> all within the 4.5-hour competition time limit Both OpenAI and Google pointed this out, but does that matter a lot? They could have spun up a million parallel reasoning processes to search for a proof that checks out - though of course some large amount of computation would have to be reserved for some kind of evaluator model to rank the proofs and decide which one to submit. Perhaps it was hundreds of years of GPU…

> They could have spun up a million parallel reasoning processes But alas, they did not, and in fact nobody did (yet). Enumerating proofs is notoriously hard for deterministic systems. I strongly recommend reading Aaronson's paper about the intersection of philosophy and complexity theory that touches these points in more detail: [1] [1]: https://www.scottaaronson.com/papers/philos.pdf

> But alas, they did not, and in fact nobody did

Seems like they actually did that:

> We achieved this year’s result using an advanced version of Gemini Deep Think – an enhanced reasoning mode for complex problems that incorporates some of our latest research techniques, including parallel thinking. This setup enables the model to simultaneously explore and combine multiple possible solutions before giving a final answer, rather than pursuing a single, linear chain of thought.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#226

Earlier quoted context omitted.

I wouldn't read too much into the timelines, as it seems that OpenAI simply broke an embargo that the other players were up to that point respecting: https://arstechnica.com/ai/2025/07/openai-jumps-gun-on-inter... Very in character for them!

Sounds like no one requested that OpenAI wait a week: >We weren't in touch with IMO. I spoke with one organizer before the post to let him know. He requested we wait until after the closing ceremony ends to respect the kids, and we did. https://x.com/polynoamial/status/1947024171860476264?s=46 https://x.com/polynoamial/status/1947398531259523481?s=46 (I work at OpenAI, but was not part of this work)

https://x.com/Mihonarium/status/1946880931723194389

OpenAI jumped the gun before *the closing party* - didn't even let the kids celebrate their winning properly before stealing the spotlight

Also in the article

> In response to the controversy, OpenAI research scientist Noam Brown posted on X, "We weren't in touch with IMO. I spoke with one organizer before the post to let him know. He requested we wait until after the closing ceremony ends to respect the kids, and we did."

> However, an IMO coordinator told X user Mikhail Samin that OpenAI actually announced before the closing ceremony, contradicting Brown's claim. The coordinator called OpenAI's actions "rude and inappropriate," noting that OpenAI "wasn't one of the AI companies that cooperated with the IMO on testing their models."

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#227

Earlier quoted context omitted.

Feynman isn't a god and he didn't conceive of AI.

Von Neumann pretty much was a god, and even as early as 1945, he was explicitly assuming that neural models were how things in the computing business would eventually shake out. The rest of us are just a little slow to catch up, is all.

Even if he did predict what would happen, it doesn't mean it's a good thing. von Neumann was a great prodigy but he was also a freak in a way, being able to enjoy the peak of very advanced technical things, which is not necessarily a good thing for the rest of us.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#228
post #154

Earlier quoted context omitted.

Lean is an interactive prover, not an automated prover. Last year a lot of human effort was required to formalise the problems in Lean before the machines could get to work. This year you get natural language input and output, and much faster. The advantage of Lean is that the system checks the solutions, so hallucination is impossible. Of course, one still relies on the problems and solutions being translated to nat…

Oh, I didn't realize that last year the problem formalization was a human effort; I assumed the provers themselves took the problem and created the formalization. Is this step actually harder to automate than solving the problem once it's formalized? Anyway mainly I was curious whether using an interactive prover like Lean would have provided any advantage, or whether that is no longer really the case. My initial tak…

I have not been working on formalization but theorem proving, so I can't confidently answer some of those questions.

However, I recognise that there is not so much training data for LLMs wanting to use the Lean language. Moreover, you are really teaching it how to apply "Lean tactics", which may or may not be related to what mathematicians do in standard texts on which LLMs have trained. Finally, the foundations are different: dependent type theory, instead of the set theory that mathematicians use.

My own personal perspective is that esoteric formal languages serve a purpose, but not this one. Most mathematicians have not been hot on the idea (with a handful of famous exceptions). But the idea seems to have gained a lot of traction anyway.

I'd personally like to see more money put into informal symbolic theorem proving tools which can very rapidly find a solution as close to natural language and the usual foundations as possible. But funding seems to be a huge issue. Academia has been bled dry, and no one has an appetite for huge multi-year projects of that kind.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#229

This is making mathematics too systematic and mechanical, and it kills the joy of it....

It's killing the joy in everything else, so why not? At least we can all wait tables and sell crap at convenience stores when it takes over our jobs.

You're giving examples of difficult and boring jobs while there will surely still be fun jobs to be done like plumbing, dangerous construction work or prostitution

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#230
post #224
post #212

Earlier quoted context omitted.

Your Nakamura example is from 2008. That's 17 years ago. The machines have improved a lot since then, hardware and software both. I've seen Nakamura beat strong-but-still-limited bots playing "anti-computer chess" but I am fairly sure he would be eaten alive if he tried it against present-day Stockfish or Leela on good hardware. Maybe you're right about correspondence chess. That interview is from 2018 and the machin…

I'm aware the Nakamura example is old, but the core issue (the horizon effect) is still there in any alpha/beta pruning engine, including the newest SF. But I will certainly grant you that it has become much harder to execute since then. MCTS engines (like Lc0) are far more immune to the horizon effect, but can instead suffer from missing very shallow tactics, especially in very fast time controls, as Andrew Tang sho…

There's a beautiful game between SF and Lc0 (a few months ago) where Stockfish thinks it's winning, while Lc0 has a lock on a draw. Lc0 then proceeds to sack 4 pieces and draw, but SF (with NNUE) only "sees" the draw 2 moves into the position.
Post reply on HN