Earlier quoted context omitted.
The trouble is, getting an IMO gold medal is much easier (by frequency) than being the #1 Go player in the world, which was achieved by AI 10 years ago. I'm not sure it's enough to just gesture at the task; drilling down into precisely how it was achieved feels important. (Not to take away from the result, which I'm really impressed by!)
The "AI" that won Go was Monte Carlo tree search on a neural net "memory" of the outcome of millions of previous games; this is a LLM solving open ended problems. The tasks are hardly even comparable.
OpenAI claims gold-medal performance at IMO 2025
141–150 of 737 posts
Re: OpenAI claims gold-medal performance at IMO 2025
#142Earlier quoted context omitted.
The key bit here is whether the LLM doing the cherry picking had knowledge of the solution. If it didn't, this is a meaningful result. That's why I'd like more info, but I fear OpenAI is going to try to keep things under wraps.
> If it didn't We kind of have to assume it didn't right? Otherwise bragging about the results makes zero sense and would be outright misleading.
why would not they? what are the incentives not to?
Re: OpenAI claims gold-medal performance at IMO 2025
#143Noam Brown: > this isn’t an IMO-specific model. It’s a reasoning LLM that incorporates new experimental general-purpose techniques. > it’s also more efficient [than o1 or o3] with its thinking. And there’s a lot of room to push the test-time compute and efficiency further. > As fast as recent AI progress has been, I fully expect the trend to continue. Importantly, I think we’re close to AI substantially contributing…
How is a claim , "clear evidence" to anything?
Re: OpenAI claims gold-medal performance at IMO 2025
#144Noam Brown: > this isn’t an IMO-specific model. It’s a reasoning LLM that incorporates new experimental general-purpose techniques. > it’s also more efficient [than o1 or o3] with its thinking. And there’s a lot of room to push the test-time compute and efficiency further. > As fast as recent AI progress has been, I fully expect the trend to continue. Importantly, I think we’re close to AI substantially contributing…
How is a claim , "clear evidence" to anything?
Unlike seemingly most here on HN, I judge people's trustworthiness individually and not solely by the organization they belong to. Noam Brown is a well known researcher in the field and I see no reason to doubt these claims other than a vague distrust of OpenAI or big tech employees generally which I reject.
Re: OpenAI claims gold-medal performance at IMO 2025
#145The cynicism/denial on HN about AI is exhausting. Half the comments are some weird form of explaining away the ever increasing performance of these models I've been reading this website for probably 15 years, its never been this bad. many threads are completely unreadable, all the actual educated takes are on X, its almost like there was a talent drain
Probably because both sides have strong vested interests and it’s next to impossible to find a dispassionate point of view. The Pro AI crowd, VC, tech CEOs etc have strong incentive to claim humans are obsolete. Many tech employees see threats to their jobs and want to poopoo any way AI could be useful or competitive.
Re: OpenAI claims gold-medal performance at IMO 2025
#146Re: OpenAI claims gold-medal performance at IMO 2025
#147Earlier quoted context omitted.
While I usually enjoy seeing these discussions, I think they are really pushing the usefulness of bayesian statistics. If one dude says the chance for an outcome is 8% and another says it's 16% and the outcome does occur, they were both pretty wrong, even though it might seem like the one who guessed a few % higher might have had a better belief system. Now if one of them had said 90% while the other said 8% or 16%,…
From a mathematical point of view there are two factors: (1) Initial prior capability of prediction from the human agents and (2) Acceleration in the predicted event. Now we examine the result under such a model and conclude that: The more prior predictive power of human agents imply the more a posterior acceleration of progress in LLMs (math capability). Here we are supposing that the increase in training data is no…
(1) Bad prior prediction capability of humans imply that result does not provide any information
(2) Good prior prediction capability of humans imply that there is acceleration in math capabilities of LLMs.
Re: OpenAI claims gold-medal performance at IMO 2025
#148This is an awesome progress in human achievement to get these machines intelligent. And this is also a fast regress and decline on the human wisdom! We are simply greasing the grooves and letting things slide faster and faster and calling it progress. How does this help to make the human and nature integration better? Does this improve climate or make humans adapt better to changing climate? Are the intelligent machi…
Nobody knows the answers to these questions. Relying on AGI solving problems like climate change seems like a risky strategy but on the other hand it’s very plausible that these tools can help in some capacity. So we have to build, study and find out but also consider any opportunity cost of building these tools versus others.
No human has any idea how to accomplish that. If a machine could, we would all have much to learn from it.
Re: OpenAI claims gold-medal performance at IMO 2025
#149Earlier quoted context omitted.
- AI competing is "wholly unfair" - "[AI is] far away from being substantially being better than MCTs" ^ pick only one
Yeah it’s a completely fair playing field, it’s completely obvious that AI should be able to compete with humans in the same way that robotics and computers can compete with humanity (and are better suited for many tasks). Whether or not they’re far away from being better than humans is up to debate, but the entire point of these types of benchmarks it to compare them to humans.
Yeah same way computers and robots should be able to win World Chess Championship, 100m dash and Wimbledon.
>>but the entire point of these types of benchmarks it to compare them to humans
The entire point of the competition is to fight against participants who are similar to you, have similar capabilities and go through similar struggles. If you want bot vs human competitions - great - organize it yourself instead of hijacking well established competitions out there.
Re: OpenAI claims gold-medal performance at IMO 2025
#150The cynicism/denial on HN about AI is exhausting. Half the comments are some weird form of explaining away the ever increasing performance of these models I've been reading this website for probably 15 years, its never been this bad. many threads are completely unreadable, all the actual educated takes are on X, its almost like there was a talent drain