Live data from Hacker News

Gemini with Deep Think achieves gold-medal standard at the IMO

deepmind.google

51–60 of 254 posts

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#51
post #43

Earlier quoted context omitted.

OpenAI announced their results after the closing ceremony as was requested. https://x.com/polynoamial/status/1947024171860476264?s=46

> as was requested. They requested week after?

He claims nobody made that request to OpenAI. It was a request made to Google and others who were being judged by the actual contest judges, which OpenAI was not.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#53

From Terence Tao, via mastodon [0]: > It is tempting to view the capability of current AI technology as a singular quantity: either a given task X is within the ability of current tools, or it is not. However, there is in fact a very wide spread in capability (several orders of magnitude) depending on what resources and assistance gives the tool, and how one reports their results. > One can illustrate this with a hum…

Discussed here:

A human metaphor for evaluating AI capability - https://news.ycombinator.com/item?id=44622973 - July 2025 (30 comments)

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#54
post #6

Advanced Gemini, not Gemini Advanced. Thanks, Google. Maybe they should have named it MathBard.

Google and Microsoft continuing to prove that the hardest problem in programming is naming things.

There are only so many words in the modern English language to hint at "next upgrade": pro, plus, ultra, new, advanced, magna, X, Z, and Ultimate. Even fewer words to explain miniaturized: mini, lite, and zero. And marketers are trying to seesaw on known words without creating new ones to explain new tech. This is why we have Bard and Gemini and Chat and Copilot.

Taking a step back, it is overused exaggeration to the point where words run out quick and newer tech needs to fight with existing words for dominance. Copilot should be the name of an AI agent. Bard should have been just a text generator. Gemini is the name of a liar. Chat is probably the iphone of naming but the GPT suffix says Creativity had not come to work that day.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#55
post #44

Earlier quoted context omitted.

> Was OpenAI simply not coordinating with the IMO Board then? You are still surprised by sama@'s asinineness? You must be new here.

When your goal is to control as much of the world's money as possible, preferably all of it, then everyone is your enemy, including high school students.

How dare those high school students use their brains to compete with ChatGPT and deny the shareholders their value?

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#57
post #8

Seems OpenAI knew this is forthcoming so they front ran the news? I was really high on Gemini 2.5 Pro after release but I kept going back to o3 for anything I cared about.

I regularly have the opposite experience: o3 is almost unusable, and Gemini 2.5 Pro is reliably great. Claude Opus 4 is a close second.

o3 is so bad it makes me wonder if I'm being served a different model? My o3 responses are so truncated and simplified as to be useless. Maybe my problems aren't a good fit, but whatever it is: o3 output isn't useful.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#59
Comparing the answers between Openai and Gemini the writing style of Gemini is a lot clearer. It could be presented a bit better but it's easy enough to follow the proof. This also makes it a lot shorter than the answer given by OpenAI and it uses proper prose.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#60
post #15

Earlier quoted context omitted.

Does the IMO reuse problems? My understanding is that new problems are submitted each year and 6 are selected for each competition. The submitted problems are then published after the IMO has concluded. How would the training data contain unpublished, newly submitted problems? Obviously the training data contained similar problems, because that's what every IMO participant already studies. It seems unlikely that they…

IMO doesn't reuse problems, but Terence Tao has a Mastodon post where he explains that the first five (of six) problems are generally ones where existing techniques can be leveraged to get to the answer. The sixth problem requires considerable originality. Notably, both Gemini and OpenAI's model didn't get the sixth problem. Still quite an achievement though.

strange statement--it's not true in general for sure (3&6 typically hardest but they certainly aren't fundamentally of a different nature to other questions) this year P6 seemed to be by far the hardest though but this posthoc statement should be read cautiously
Post reply on HN