Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

321–330 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#321

Wow. That's an impressive result, but how did they do it? Wei references scaling up test-time compute, so I have to assume they threw a boatload of money at this. I've heard talk of running models in parallel and comparing results - if OpenAI ran this 10000 times in parallel and cherry-picked the best one, this is a lot less exciting. If this is legit, then we need to know what tools were used and how the model used…

Why is that less exciting? A machine competing in an unconstrained natural language difficult math contest and coming out on top by any means is breath taking science fiction a few years ago - now it’s not exciting? Regardless of the tools for verification or even solvers - why is the goal post moving so fast? There is no bonus for “purity of essence” and using only neural networks. We live in an era where it’s hard…

Without sharing their methodology, how can we trust the claim ? questions like:

1) did humans formalize the input 2) did humans prompt the llm towards the solution etc..

I am excited to hear about it, but I remain skeptical.

Re: OpenAI claims gold-medal performance at IMO 2025

#322

Earlier quoted context omitted.

> high school/early university maths problems should not have been a stretch at all for it This is a ridiculous understatement of the difficulty of getting gold at the IMO.

[flagged]

"Please don't post shallow dismissals, especially of other people's work. A good critical comment teaches us something."

You've broken this guideline repeatedly in this thread. Can you please not do that on HN?

https://news.ycombinator.com/newsguidelines.html

Re: OpenAI claims gold-medal performance at IMO 2025

#323

Earlier quoted context omitted.

> These models cannot reason Not trying to be a smarty pants here, but what do we mean by "reason"? Just to make the point, I'm using Claude to help me code right now. In between prompts, I read HN. It does things for me such as coding up new features, looking at the compile and runtime responses, and then correcting the code. All while I sit here and write with you on HN. It gives me feedback like "lock free message…

You raise a far point. These criticisms based on "it's merely X" or "it's not really Y" don't hold water when X and Y are poorly defined. The only thing that should matter is the results they get. And I have a hard time understanding why the thing that is supposed to behave in an intelligent way but often just spew nonsense gets 10x budget increases over and over again. This is bad software. It does not do the thing…

We already have highly advanced deterministic software. The value lies in the abductive “reasoning” and natural language processing.

We deal with non determinism any time our code interacts with the natural world. We build guard rails, detection, classification of false/true positive and negatives, and all that all the time. This isn’t a flaw, it’s just the way things are for certain classes of problems and solutions.

It’s not bad software - it’s software that does things we’ve been trying to do for nearly a hundred years beyond any reasonable expectation. The fact I can tell a machine in human language to do some relative abstract and complex task and it pretty reliably “understands” me and my intent, “understands” it’s tools and capabilities, and “reasons” how to bridge my words to a real world action is not bad software. It’s science fiction.

The fact “reliably” shows up is the non determinism. Not perfectly, although on a retry with a new seed it often succeeds. This feels like most software that interacts with natural processes in any way or form.

It’s remarkable that anyone who has ever implemented exponential back off and retry, has ever implemented edge cases, and sir and say “nothing else matters,” when they make their living dealing with non determinism. Because the algorithmic kernel of logic is 1% of programming and systems engineering, and 99% is coping with the non determinism in computing systems.

The technology is immature and the toolchains are almost farcically basic - money is dumping into model training because we have not yet hit a wall with brute force. And it takes longer to build a new way of programming and designing highly reliable systems in the face of non determinism, but it’s getting better faster than almost any technology change in my 35 years in the industry.

Your statement that it “very often produces wrong or nonsensical output” also tells me you’re holding onto a bias from prior experiences. The rate of improvement is astonishing. At this point in my professional use of frontier LLMs and techniques they are exceeding the precision and recall of humans and there’s a lot of rich ground untouched. At this point we largely can offload massive amounts of work that humans would do in decision making (classification) and use humans as a last line to exercise executive judgement often with the assistance of LLMs. I expect within two years humans will only be needed in the most exceptional of situations, and we will do a better job on more tasks than we ever could have dreamed of with humans. For the company I’m at this is a huge bottom line improvement far and beyond the cost of our AI infrastructure and development, and we do quite a lot of that too.

If you’re not seeing it yet, I wouldn’t use that to extrapolate to the world at large and especially not to the future.

Re: OpenAI claims gold-medal performance at IMO 2025

#324
It’s interesting that this is a competition elite enough that several posters on a programming website don’t seem to understand what it is.

My very rough napkin math suggests that against the US reference class, imo gold is literally a one in a million talent (very roughly 20 people who make camp could get gold out of very roughly twenty million relevant high schoolers).

Re: OpenAI claims gold-medal performance at IMO 2025

#325

Earlier quoted context omitted.

Which would be impressive if we knew those problems weren't in the training data already. I mean it is quite impressive how language models are able to mobilize the knowledge they have been trained on, especially since they are able to retrieve information from sources that may be formatted very differently, with completely different problem statement sentences, different variable names and so on, and really operate…

Even if we accept as a premise that these models are doing "smart retrieval" and not "reasoning" (neither of which are being defined here, nor do I think we can tell from this tweet even if they were), it doesn't really change the impact. There are many industries for which the vast majority of work done is closer to what I think you mean by "smart retrieval" than what I think you mean by "reasoning." Adult primary c…

> There are many industries for which the vast majority of work done is closer to what I think you mean by "smart retrieval" than what I think you mean by "reasoning." Adult primary care and pediatrics, finance, law, veterinary medicine, software engineering, etc. At least half, if not upwards of 80% of the work in each of these fields is effectively pattern matching to a known set of protocols. They absolutely deal in novel problems as well, but it's not the majority of their work.

I wholeheartedly agree with that.

I'm in fact pretty bullish on LLMs, as tools with near infinite industrial use cases, but I really dislike the “AGI soon” narrative (which sets expectations way too high).

IMHO the biggest issue with LLMs isn't that they aren't good enough at solving math problem, but that there's no easy way to add information to a model after its training, which is a significant problem for a “smart information retrieval” system. RAG is used as a hack around this issue, but its performance can vary a ton with tasks. LORAs are another options, but they require significant work to make a dataset, and you can only cross your fingers the model keeps its abilities.

Re: OpenAI claims gold-medal performance at IMO 2025

#327
post #186

The AI scaling that went on for the last five years is going to be very different from the scaling that will happen in the next ten years. These models have latent capabilities that we are racing to unearth. IMO is but one example. There’s so much to do at inference time. This result could not have been achieved without the substrate of general models. Its not like Go or protein folding. You need the collective publi…

Latent?? If you looked at RLHF hiring over the last year, there was a huge hiring of IMO competitors to RLHF. This was a new, highly targeted, highly funded RLHF’ing.

Can you provide any kind of source? Very curious about this!

Re: OpenAI claims gold-medal performance at IMO 2025

#328
post #324

It’s interesting that this is a competition elite enough that several posters on a programming website don’t seem to understand what it is. My very rough napkin math suggests that against the US reference class, imo gold is literally a one in a million talent (very roughly 20 people who make camp could get gold out of very roughly twenty million relevant high schoolers).

I’m not trying to take away from the difficulty of the competition. But I went to a relatively well regarded high school and never even heard of IMO until I met competitors during undergrad.

I think that the number of students who are even aware of the competition is way lower than the total number of students.

I mean, I don’t think I’d have been a great competitor even if I tried. But I’m pretty sure there are a lot of students that could do well if given the opportunity.

Re: OpenAI claims gold-medal performance at IMO 2025

#329
post #44

From that thread: "The model solved P1 through P5; it did not produce a solution for P6." It's interesting that it didn't solve the problem that was by far the hardest for humans too. China, the #1 team got only 21/42 points on it. In most other teams nobody solved it.

To me, this is a tell of human-involvement in the model solution.

There is no reason why machines would do badly on exactly the problem which humans do badly as well - without humans prodding the machine towards a solution.

Also, there is no reason why machines could not produce a partial or wrong answer to problem 6 which seems like survivor bias to me. ie, that only correct solutions were cherrypicked.

Re: OpenAI claims gold-medal performance at IMO 2025

#330

OpenAI simply can’t be trusted on any benchmarks: https://news.ycombinator.com/item?id=42761648

Somewhat related, but I’ve been feeling as of late what can best be described as “benchmark fatigue”.

The latest models can score something like 70% on SWE-bench verified and yet it’s difficult to say what tangible impact this has on actual software development. Likewise, they absolutely crush humans at sport programming but are unreliable software engineers on their own.

What does it really mean that an LLM got gold on this year’s IMO? What if it means pretty much nothing at all besides the simple fact that this LLM is very, very good at IMO style problems?

Post reply on HN