Earlier quoted context omitted.
OpenAI announced their results after the closing ceremony as was requested. https://x.com/polynoamial/status/1947024171860476264?s=46
> as was requested. They requested week after?
Gemini with Deep Think achieves gold-medal standard at the IMO
51–60 of 254 posts
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#52Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#53From Terence Tao, via mastodon [0]: > It is tempting to view the capability of current AI technology as a singular quantity: either a given task X is within the ability of current tools, or it is not. However, there is in fact a very wide spread in capability (several orders of magnitude) depending on what resources and assistance gives the tool, and how one reports their results. > One can illustrate this with a hum…
A human metaphor for evaluating AI capability - https://news.ycombinator.com/item?id=44622973 - July 2025 (30 comments)
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#54Advanced Gemini, not Gemini Advanced. Thanks, Google. Maybe they should have named it MathBard.
Google and Microsoft continuing to prove that the hardest problem in programming is naming things.
Taking a step back, it is overused exaggeration to the point where words run out quick and newer tech needs to fight with existing words for dominance. Copilot should be the name of an AI agent. Bard should have been just a text generator. Gemini is the name of a liar. Chat is probably the iphone of naming but the GPT suffix says Creativity had not come to work that day.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#55Earlier quoted context omitted.
> Was OpenAI simply not coordinating with the IMO Board then? You are still surprised by sama@'s asinineness? You must be new here.
When your goal is to control as much of the world's money as possible, preferably all of it, then everyone is your enemy, including high school students.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#56Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#57Seems OpenAI knew this is forthcoming so they front ran the news? I was really high on Gemini 2.5 Pro after release but I kept going back to o3 for anything I cared about.
o3 is so bad it makes me wonder if I'm being served a different model? My o3 responses are so truncated and simplified as to be useless. Maybe my problems aren't a good fit, but whatever it is: o3 output isn't useful.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#58Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#59Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#60Earlier quoted context omitted.
Does the IMO reuse problems? My understanding is that new problems are submitted each year and 6 are selected for each competition. The submitted problems are then published after the IMO has concluded. How would the training data contain unpublished, newly submitted problems? Obviously the training data contained similar problems, because that's what every IMO participant already studies. It seems unlikely that they…
IMO doesn't reuse problems, but Terence Tao has a Mastodon post where he explains that the first five (of six) problems are generally ones where existing techniques can be leveraged to get to the answer. The sixth problem requires considerable originality. Notably, both Gemini and OpenAI's model didn't get the sixth problem. Still quite an achievement though.