Live data from Hacker News

Gemini with Deep Think achieves gold-medal standard at the IMO

deepmind.google

211–220 of 254 posts

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#211

Earlier quoted context omitted.

Gemini is clearer but MY GOD is it verbose. e.g. look at problem 1, section 2. Analysis of the Core Problem - there's nothing at all deep here, but it seems the model wants to spell out every single tiny logical step. I wonder if this is a stylistic choice or something that actually helps the model get to the end.

They actually do help - in that they give the model more computation time and also allow realtime management of the input context by the model. You can see this same behavior in the excessive comment writing some coding models engage in; Anthropic interviews said these do actually help the model.

Gemini did not one-shot these answers; it did its thinking elsewhere (probably not released by Google) and then it consolidated it down into what you see in the PDF. From the article:

> We achieved this year’s result using an advanced version of Gemini Deep Think – an enhanced reasoning mode for complex problems that incorporates some of our latest research techniques, including parallel thinking. This setup enables the model to simultaneously explore and combine multiple possible solutions before giving a final answer, rather than pursuing a single, linear chain of thought.

I don't see any parallel thinking, e.g., so that was probably elided in the final results.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#212
post #207
post #173

Earlier quoted context omitted.

My understanding is that they did , but don't any more; it's no longer true that humans understand enough things about chess better than computers for the human/computer collaboration to contribute anything over just using the computer. I don't think the interval between "computers are almost as strong as humans" and "computers are so much stronger than humans that there's no way for even the strongest humans to cont…

> My understanding is that they did, but don't any more; it's no longer true that humans understand enough things about chess better than computers for the human/computer collaboration to contribute anything over just using the computer. This is not true, at least not in very long time formats like correspondence chess: https://en.chessbase.com/post/correspondence-chess-and-corre... There's also many well known cases…

Your Nakamura example is from 2008. That's 17 years ago. The machines have improved a lot since then, hardware and software both. I've seen Nakamura beat strong-but-still-limited bots playing "anti-computer chess" but I am fairly sure he would be eaten alive if he tried it against present-day Stockfish or Leela on good hardware.

Maybe you're right about correspondence chess. That interview is from 2018 and the machines have got distinctly stronger in that time, but 7 years isn't so long and it could be that human input still has some value for CC.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#214

Earlier quoted context omitted.

If a Language Model is capable of producing rigorous natural language proofs then getting it to produce Lean (or whatever) proofs would not be a big deal. Lean use in AlphaProof was something of a crutch (not saying this as a bad thing). Very specialized, very narrow with little use outside any other domain. On the other hand, if you can achieve the same with general RL techniques and natural language then other hard…

If a Language Model is capable of producing rigorous natural language proofs then getting it to produce Lean (or whatever) proofs would not be a big deal. This is a wildly uninformed take. Even today there are plenty of basic statements which LLM’s can produce English language proofs of that have not been formalized.

Most mathematicians aren't much interested in translating statements for the fun of it, so whether a lot of basic statements are un-formalized doesn't mean much. And the point was never that formalization was easy.

I said that if language models become capable enough of not needing a crutch, then adding one afterwards isn't a big deal. What exactly do you think Alphaproof is? Much worse LLMs were already doing what you are saying. There's a reason it preceded this and not the other way around.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#215

This year, our advanced Gemini model operated end-to-end in natural language, producing rigorous mathematical proofs directly from the official problem descriptions I think I have a minority opinion here, but I’m a bit disappointed they seem to be moving away from formal techniques. I think if you ever want to truly “automate” math or do it at machine scale, e.g. creating proofs that would amount to thousands of page…

(Stream of consciousness aside: That said, letting machines go wild in the depths of the consequences of some axiomatic system like ZFC may reveal a method of proof mathematicians would find to be monstrous. So like, if ZFC is inconsistent, then anything can be proven. But short of that, maybe the machines will find extremely powerful techniques which “almost” prove inconsistency that nevertheless somehow lead to log…

The horror scenario you describe would actually be more valuable than the insinuated "spirit" I believe:

Suppose it faithfully reasons and attempts to find proofs of claims, in the best case you found a proof of a specific claim (IN AN INCONSISTENT SYSTEM).

Suppose in the "horror scenario" that the machine has surreptitiously found a proof of false in ZFC (and can now prove any claim), and is not disclosing it, but abusing it to present 'actual proofs in inconsistent ZFC' for whatever claims the user asks it. In this case we can just ask for a proof of A and a proof of !A, if it proves both it has leaked the fact it found and exploits an inconsistency in the formal system! Thats worth more than a hard to find proof, in an otherwise inconsistent system.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#216
post #211

Earlier quoted context omitted.

They actually do help - in that they give the model more computation time and also allow realtime management of the input context by the model. You can see this same behavior in the excessive comment writing some coding models engage in; Anthropic interviews said these do actually help the model.

Gemini did not one-shot these answers; it did its thinking elsewhere (probably not released by Google) and then it consolidated it down into what you see in the PDF. From the article: > We achieved this year’s result using an advanced version of Gemini Deep Think – an enhanced reasoning mode for complex problems that incorporates some of our latest research techniques, including parallel thinking. This setup enables…

Yes, because these are the answers it gave, not the thinking.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#219
post #191
post #166

Earlier quoted context omitted.

>> > I think the reason why PG respects Sam so much is he is charismatic, resourceful, and just overall seems like a genuine person. does he? wasn't sama ousted of YC in some muddy ways after he tried to co-opt in into an OpenAI investment arm, was funny to find the YC Open Research project landing page on yc's website now defunct and pointing how he misrepresented it as a YC project when it was his own maybe he fear…

The post was from 14 years ago, before that.

oh, gotcha

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#220
I find this damn impressive, but I’m disappointed that while there’s some version of the final proofs released, I haven’t seen the reasoning traces. Although, I’ll admit the only reason I want to see them is because I hope they shed some light on how the models improved so much so quickly. I’d also like to see confirmation of how much compute was used.
Post reply on HN