Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

81–90 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#81

These are high school level only in the sense of assumed background knowledge, they are extremely difficult. Professional mathematicians would not get this level of performance, unless they have a background in IMO themselves. This doesn’t mean that the model is better than them in math, just that mathematicians specialize in extending the frontier of math. The answers are not in the training data. This is not a mode…

Are you sure this is not specialized to IMO? I do see the twitter thread saying it's "general reasoning" but I'd imagine they RL'd on olympiad math questions? If not I really hope someone from OpenAI says that bc it would be pretty astounding.

They also said this is not part of GPT-5, and “will be released later”. It’s very, very likely a model specifically fine-tuned for this benchmark, where afterwards they’ll evaluate what actual real-world problems it’s good at (eg like “use o4-mini-high for coding”).

Re: OpenAI claims gold-medal performance at IMO 2025

#82
post #63
post #7

Earlier quoted context omitted.

[flagged]

I feel like I've noticed you you making the same comment 12 places in this thread -- incorrectly misrepresenting the difficulty of this tournament and ultimately it comes across as a bitter ex. Here's an example problem 5: Let a1,a2,…,an be distinct positive integers and let M=max⁡1≤i Find the maximum number of pairs (i,j) with 1≤i<j≤n for which (ai +aj )(aj −ai )=M.

What does max⁡1≤i<j≤n mean? Wouldn't M always be j?

Re: OpenAI claims gold-medal performance at IMO 2025

#83
post #70

Earlier quoted context omitted.

Impressive prediction, especially pre-ChatGPT. Compare to Gary Marcus 3 months ago: https://garymarcus.substack.com/p/reports-of-llms-mastering-... We may certainly hope Eliezer's other predictions don't prove so well-calibrated.

I do think Gary Marcus says a lot of wrong stuff about LLMs but I don’t see anything too egregious in that post. He’s just describing the results they got a few months ago.

He definitely cannot use the original arguments from then ChatGPT arrived, he's a perennial goal post shifter.

Re: OpenAI claims gold-medal performance at IMO 2025

#84

These are high school level only in the sense of assumed background knowledge, they are extremely difficult. Professional mathematicians would not get this level of performance, unless they have a background in IMO themselves. This doesn’t mean that the model is better than them in math, just that mathematicians specialize in extending the frontier of math. The answers are not in the training data. This is not a mode…

It almost certainly is specialized to IMO problems, look at the way it is answering the questions: https://xcancel.com/alexwei_/status/1946477742855532918 E.g here: https://pbs.twimg.com/media/GwLtrPeWIAUMDYI.png?name=orig Frankly it looks to me like it's using an AlphaProof style system, going between natural language and Lean/etc. Of course OpenAI will not tell us any of this.

I actually think this “cheating” is fine. In fact it’s preferable. I don’t need an AI that can act as a really expensive calculator or solver. We’ve already built really good calculators and solvers that are near optimal. What has been missing is the abductive ability to successfully use those tools in an unconstrained space with agency. I find really no value in avoiding the optimal or near optimal techniques we’ve devised rather than focusing on the harder reasoning tasks of choosing tools, instrumenting them properly, interpreting their results, and iterating. This is the missing piece in automated reasoning after all. A NN that can approximate at great cost those tools is a parlor trick and while interesting not useful or practical. Even if they have some agent system here, it doesn’t make the achievement any less that a machine can zero shot do as well as top humans at incredibly difficult reasoning problems posed in natural language.

Re: OpenAI claims gold-medal performance at IMO 2025

#86

Earlier quoted context omitted.

>if OpenAI ran this 10000 times in parallel and cherry-picked the best one, this is a lot less exciting. That entirely depends on who did the cherry picking. If the LLM had 10000 attempts and each time a human had to falsify it, this story means absolutely nothing. If the LLM itself did the cherry picking, then this is just akin to a human solving a hard problem. Attempting solutions and falsifying them until the des…

The key bit here is whether the LLM doing the cherry picking had knowledge of the solution. If it didn't, this is a meaningful result. That's why I'd like more info, but I fear OpenAI is going to try to keep things under wraps.

> If it didn't

We kind of have to assume it didn't right? Otherwise bragging about the results makes zero sense and would be outright misleading.

Re: OpenAI claims gold-medal performance at IMO 2025

#87
I am neither an optimist nor a pessimist for AI. I would likely be called both by the opposite parties. But the fact that AI / LLM is still rapidly improving is impressive in itself and worth celebrating for. Is it perfect, AGI, ASI? No. Is it useless? Absolutely not.

I am just happy the prize is so big for AI that there are enough money involve to push for all the hardware advancement. Foundry, Packaging, Interconnect, Network etc, all the hardware research and tech improvements previously thought were too expensive are now in the "Shut up and take my money" scenario.

Re: OpenAI claims gold-medal performance at IMO 2025

#89
Am I missing something or is this completely meaningless? It's 100% opaque, no details whatsoever and no transparency or reproducibility.

I wouldn't trust these results as it is. Considering that there are trillions of dollars on the line as a reward for hyping up LLMs, I trust it even less.

Re: OpenAI claims gold-medal performance at IMO 2025

#90
post #88

OpenAI simply can’t be trusted on any benchmarks: https://news.ycombinator.com/item?id=42761648

This is not a benchmark, really. It's an official test.

And what were the methods? How was the evaluation? They could be making it all up for all we know!
Post reply on HN