Live data from Hacker News

Gemini with Deep Think achieves gold-medal standard at the IMO

deepmind.google

21–30 of 254 posts

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#21

From Terence Tao, via mastodon [0]: > It is tempting to view the capability of current AI technology as a singular quantity: either a given task X is within the ability of current tools, or it is not. However, there is in fact a very wide spread in capability (several orders of magnitude) depending on what resources and assistance gives the tool, and how one reports their results. > One can illustrate this with a hum…

Unlike OpenAI, Deepmind at least signed up for the competition ahead of time.

Agree with Tao though, I am skeptical of any result of this type unless there's a lot of transparency, ideally ahead of time. If not ahead of time, then at least the entire prompt and fine-tune data that was used.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#22
post #8

Seems OpenAI knew this is forthcoming so they front ran the news? I was really high on Gemini 2.5 Pro after release but I kept going back to o3 for anything I cared about.

>I was really high on Gemini 2.5 Pro after release but I kept going back to o3 for anything I cared about

Same here. I was impressed by their benchmarks and topping most leaderboards, but in day to day use they still feel so far behind.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#23

Do I understand it correctly that OpenAI self-proclaimed that they got their gold, without the official IMO judges grading their solutions?

Childish. And, of course they must have known there was an official LLM cohort taking the real test, and they probably even knew that Gemini got a gold medal, and may have even known that Google planned a press release for today.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#24
Related news:

- OpenAI claims gold-medal performance at IMO 2025 https://news.ycombinator.com/item?id=44613840

- "According to a friend, the IMO asked AI companies not to steal the spotlight from kids and to wait a week after the closing ceremony to announce results. OpenAI announced the results BEFORE the closing ceremony.

According to a Coordinator on Problem 6, the one problem OpenAI couldn't solve, "the general sense of the IMO Jury and Coordinators is that it was rude and inappropriate" for OpenAI to do this.

OpenAI wasn't one of the AI companies that cooperated with the IMO on testing their models, so unlike the likely upcoming Google DeepMind results, we can't even be sure OpenAI's "gold medal" is legit. Still, the IMO organizers directly asked OpenAI not to announce their results immediately after the olympiad.

Sadly, OpenAI desires hype and clout a lot more than it cares about letting these incredibly smart kids celebrate their achievement, and so they announced the results yesterday." https://x.com/mihonarium/status/1946880931723194389

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#25
> Btw as an aside, we didn’t announce on Friday because we respected the IMO Board's original request that all AI labs share their results only after the official results had been verified by independent experts & the students had rightly received the acclamation they deserved

> We've now been given permission to share our results and are pleased to have been part of the inaugural cohort to have our model results officially graded and certified by IMO coordinators and experts, receiving the first official gold-level performance grading for an AI system!

From https://x.com/demishassabis/status/1947337620226240803

Was OpenAI simply not coordinating with the IMO Board then?

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#27
post #8

Seems OpenAI knew this is forthcoming so they front ran the news? I was really high on Gemini 2.5 Pro after release but I kept going back to o3 for anything I cared about.

>I was really high on Gemini 2.5 Pro after release but I kept going back to o3 for anything I cared about Same here. I was impressed by their benchmarks and topping most leaderboards, but in day to day use they still feel so far behind.

I think that's most likely just your view, and not really based on evidence.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#28

Do I understand it correctly that OpenAI self-proclaimed that they got their gold, without the official IMO judges grading their solutions?

(deleted because I was mistaken)

isn't he an IOI medalist? and even if he was an IMO medalist, isn't there a bit of a conflict of interests?

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#29

Do I understand it correctly that OpenAI self-proclaimed that they got their gold, without the official IMO judges grading their solutions?

(deleted because I was mistaken)

I'm pretty sure when they got the gold medal they weren't allowed to judge themselves.
Post reply on HN