Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

291–300 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#291

Earlier quoted context omitted.

Based on the past history with frontier-math & AIME 2025 [1],[2] I would not trust announcements which cant be independently verified. I am excited to try it out though. Also, the performance of LLMs on imo 2025 was not even bronze [3]. Finally, this article shows that LLMs were just mostly bluffing [4] on usamo 2025. [1] https://www.reddit.com/r/slatestarcodex/comments/1i53ih7/fro ... [2] https://x.com/DimitrisPapai…

The solutions were publicly posted to GitHub: https://github.com/aw31/openai-imo-2025-proofs/tree/main

Did humans formalize the inputs ? or was the exact natural language input provided to the llm. A lot of detail is missing on the methodology used. Not to mention of any independent validation.

My skepticism stems from the past frontier math announcement which turned out to be a bluff.

Re: OpenAI claims gold-medal performance at IMO 2025

#292

I think equally impressive is the performance of the OpenAI team at the "AtCoder World Tour Finals 2025" a couple of days ago. There were 12 human participants and only one did better than OpenAI. Not sure there is a good writeup about it yet but here is the livestream: https://www.youtube.com/live/TG3ChQH61vE .

And yet when working on production code current LLMs are about as good as a poor intern. Not sure why the disconnect.

because competitive coding is narrow well described domain(limited number of concepts: lists, trees, etc) with high volume of data available for training, and easy way to setup RL feeback loop, so models can improve well in this domain, which is not true about typical enterprise overbloated software.

Re: OpenAI claims gold-medal performance at IMO 2025

#293
post #86

Earlier quoted context omitted.

> If it didn't We kind of have to assume it didn't right? Otherwise bragging about the results makes zero sense and would be outright misleading.

openai have been caught doing exactly this before

Why do people keep making up controversial claims like this? There is no evidence at all to this effect

Re: OpenAI claims gold-medal performance at IMO 2025

#295

These are high school level only in the sense of assumed background knowledge, they are extremely difficult. Professional mathematicians would not get this level of performance, unless they have a background in IMO themselves. This doesn’t mean that the model is better than them in math, just that mathematicians specialize in extending the frontier of math. The answers are not in the training data. This is not a mode…

It almost certainly is specialized to IMO problems, look at the way it is answering the questions: https://xcancel.com/alexwei_/status/1946477742855532918 E.g here: https://pbs.twimg.com/media/GwLtrPeWIAUMDYI.png?name=orig Frankly it looks to me like it's using an AlphaProof style system, going between natural language and Lean/etc. Of course OpenAI will not tell us any of this.

OpenAI explicitly stated that it is natural language only, with no tools such as Lean.

https://x.com/alexwei_/status/1946477745627934979?s=46&t=Hov...

Re: OpenAI claims gold-medal performance at IMO 2025

#296

Earlier quoted context omitted.

We can only go off their word unfortunately and they say no formal math. so I assume it's being eval'd by a verifier model instead of a formal system. There's actually some hints of this b/c geometry in Lean is not that well developed so unless they also built their own system it's hard to do it formally (though their P2 proof is by coordinate bash (computation by algebra instead of geometric construction) so it's ha…

> We can only go off their word We’re talking about Sam Altman’s company here. The same company that started out as a non profit claiming they wanted to better the world. Suggesting they should be given the benefit of the doubt is dishonest at this point.

“they must be lying because I personally dislike them”

This is why HN threads about AI have become exhausting to read

Re: OpenAI claims gold-medal performance at IMO 2025

#297

Noam Brown: > this isn’t an IMO-specific model. It’s a reasoning LLM that incorporates new experimental general-purpose techniques. > it’s also more efficient [than o1 or o3] with its thinking. And there’s a lot of room to push the test-time compute and efficiency further. > As fast as recent AI progress has been, I fully expect the trend to continue. Importantly, I think we’re close to AI substantially contributing…

How is a claim , "clear evidence" to anything?

I read the GP's comment as "but [assuming this claim is correct], this is clear evidence to the contrary."

Re: OpenAI claims gold-medal performance at IMO 2025

#298
post #180

Earlier quoted context omitted.

The International Math Olympiad isn’t an AI benchmark. It’s an annual human competition.

They didn’t actually compete.

Correct, they took the problems from the competition.

It’s not an AI benchmark generated for AI. It was targeted at humans

Re: OpenAI claims gold-medal performance at IMO 2025

#299

The cynicism/denial on HN about AI is exhausting. Half the comments are some weird form of explaining away the ever increasing performance of these models I've been reading this website for probably 15 years, its never been this bad. many threads are completely unreadable, all the actual educated takes are on X, its almost like there was a talent drain

Two things can happen at the same time: Genuine technological progress and the “hype machine” going into absolute overdrive.

The problem with the hype machine is that it provokes an opposite reaction and the noise from it buries any reasonable / technical discussion.

Re: OpenAI claims gold-medal performance at IMO 2025

#300

Earlier quoted context omitted.

The solutions were publicly posted to GitHub: https://github.com/aw31/openai-imo-2025-proofs/tree/main

Did humans formalize the inputs ? or was the exact natural language input provided to the llm. A lot of detail is missing on the methodology used. Not to mention of any independent validation. My skepticism stems from the past frontier math announcement which turned out to be a bluff.

People are reading a lot into the FrontierMath articles from a couple months ago, but tbh I don’t really understand what the controversy is supposed to be there. failing to clearly disclose sponsoring Epoch to make the benchmark clearly doesn’t affect performance of a model on it
Post reply on HN