From Terence Tao, via mastodon [0]: > It is tempting to view the capability of current AI technology as a singular quantity: either a given task X is within the ability of current tools, or it is not. However, there is in fact a very wide spread in capability (several orders of magnitude) depending on what resources and assistance gives the tool, and how one reports their results. > One can illustrate this with a hum…
Gemini with Deep Think achieves gold-medal standard at the IMO
41–50 of 254 posts
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#42Still no information on the amount of compute needed; would be interested to see a breakdown from Google or OpenAI on what it took to achieve this feat. Something that was hotly debated in the thread with OpenAI's results: "We also provided Gemini with access to a curated corpus of high-quality solutions to mathematics problems, and added some general hints and tips on how to approach IMO problems to its instructions…
Ok but when reported by mass media, which never used SI units and instead uses units like libraries of Congress, or elephants, what kind of unit should media use to compare computational energy of ai vs children?
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#43Related news: - OpenAI claims gold-medal performance at IMO 2025 https://news.ycombinator.com/item?id=44613840 - "According to a friend, the IMO asked AI companies not to steal the spotlight from kids and to wait a week after the closing ceremony to announce results. OpenAI announced the results BEFORE the closing ceremony. According to a Coordinator on Problem 6, the one problem OpenAI couldn't solve, "the general s…
OpenAI announced their results after the closing ceremony as was requested. https://x.com/polynoamial/status/1947024171860476264?s=46
They requested week after?
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#44> Btw as an aside, we didn’t announce on Friday because we respected the IMO Board's original request that all AI labs share their results only after the official results had been verified by independent experts & the students had rightly received the acclamation they deserved > We've now been given permission to share our results and are pleased to have been part of the inaugural cohort to have our model results off…
> Was OpenAI simply not coordinating with the IMO Board then? You are still surprised by sama@'s asinineness? You must be new here.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#45So, the problem wasn't translated to Lean first. But did the model use Lean, or internet search, or a calculator or Python or any other tool during its internal thinking process? OpenAI said theirs didn't, and I'm not sure if this is exactly the same claim. More clarity on this point would be nice.
I would also love to know the rough order of magnitude of the amount of computation used by both systems, measured in dollars. Being able to do it at all is of course impressive, but not useful yet if the price is outrageous. In the absence of disclosure I'm going to assume the price is, in fact, outrageous.
Edit: "No tool use, no internet access" confirmed: https://x.com/FredZhang0/status/1947364744412758305
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#46Still no information on the amount of compute needed; would be interested to see a breakdown from Google or OpenAI on what it took to achieve this feat. Something that was hotly debated in the thread with OpenAI's results: "We also provided Gemini with access to a curated corpus of high-quality solutions to mathematics problems, and added some general hints and tips on how to approach IMO problems to its instructions…
Ok but when reported by mass media, which never used SI units and instead uses units like libraries of Congress, or elephants, what kind of unit should media use to compare computational energy of ai vs children?
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#47From Terence Tao, via mastodon [0]: > It is tempting to view the capability of current AI technology as a singular quantity: either a given task X is within the ability of current tools, or it is not. However, there is in fact a very wide spread in capability (several orders of magnitude) depending on what resources and assistance gives the tool, and how one reports their results. > One can illustrate this with a hum…
I don't think anybody thinks AI was competing fair and within the rules that apply to humans. But if the humans were competing on the terms that AI solved those problems on, near-unlimited access to energy, raw compute and data, still very few humans could solve those problems within a reasonable timeframe. It would take me probably months or years to educate myself sufficiently to even have a chance.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#48Related news: - OpenAI claims gold-medal performance at IMO 2025 https://news.ycombinator.com/item?id=44613840 - "According to a friend, the IMO asked AI companies not to steal the spotlight from kids and to wait a week after the closing ceremony to announce results. OpenAI announced the results BEFORE the closing ceremony. According to a Coordinator on Problem 6, the one problem OpenAI couldn't solve, "the general s…
OpenAI announced their results after the closing ceremony as was requested. https://x.com/polynoamial/status/1947024171860476264?s=46
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#49Still no information on the amount of compute needed; would be interested to see a breakdown from Google or OpenAI on what it took to achieve this feat. Something that was hotly debated in the thread with OpenAI's results: "We also provided Gemini with access to a curated corpus of high-quality solutions to mathematics problems, and added some general hints and tips on how to approach IMO problems to its instructions…
Ok but when reported by mass media, which never used SI units and instead uses units like libraries of Congress, or elephants, what kind of unit should media use to compare computational energy of ai vs children?
So funnily enough, "the AI wrote x times the library of Congress to get there" is good enough of a comparison.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#50Related news: - OpenAI claims gold-medal performance at IMO 2025 https://news.ycombinator.com/item?id=44613840 - "According to a friend, the IMO asked AI companies not to steal the spotlight from kids and to wait a week after the closing ceremony to announce results. OpenAI announced the results BEFORE the closing ceremony. According to a Coordinator on Problem 6, the one problem OpenAI couldn't solve, "the general s…