Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

141–150 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#141

Earlier quoted context omitted.

The trouble is, getting an IMO gold medal is much easier (by frequency) than being the #1 Go player in the world, which was achieved by AI 10 years ago. I'm not sure it's enough to just gesture at the task; drilling down into precisely how it was achieved feels important. (Not to take away from the result, which I'm really impressed by!)

The "AI" that won Go was Monte Carlo tree search on a neural net "memory" of the outcome of millions of previous games; this is a LLM solving open ended problems. The tasks are hardly even comparable.

A "reasoning LLM" might not be conceptually far from MCTS.

Re: OpenAI claims gold-medal performance at IMO 2025

#142
post #86

Earlier quoted context omitted.

The key bit here is whether the LLM doing the cherry picking had knowledge of the solution. If it didn't, this is a meaningful result. That's why I'd like more info, but I fear OpenAI is going to try to keep things under wraps.

> If it didn't We kind of have to assume it didn't right? Otherwise bragging about the results makes zero sense and would be outright misleading.

> would be outright misleading

why would not they? what are the incentives not to?

Re: OpenAI claims gold-medal performance at IMO 2025

#143

Noam Brown: > this isn’t an IMO-specific model. It’s a reasoning LLM that incorporates new experimental general-purpose techniques. > it’s also more efficient [than o1 or o3] with its thinking. And there’s a lot of room to push the test-time compute and efficiency further. > As fast as recent AI progress has been, I fully expect the trend to continue. Importantly, I think we’re close to AI substantially contributing…

How is a claim , "clear evidence" to anything?

[flagged]

Re: OpenAI claims gold-medal performance at IMO 2025

#144

Noam Brown: > this isn’t an IMO-specific model. It’s a reasoning LLM that incorporates new experimental general-purpose techniques. > it’s also more efficient [than o1 or o3] with its thinking. And there’s a lot of room to push the test-time compute and efficiency further. > As fast as recent AI progress has been, I fully expect the trend to continue. Importantly, I think we’re close to AI substantially contributing…

How is a claim , "clear evidence" to anything?

Most evidence you have about the world is claims from other people, not direct experiment. There seems to be a thought-terminating cliche here on HN, dismissing any claim from employees of large tech companies.

Unlike seemingly most here on HN, I judge people's trustworthiness individually and not solely by the organization they belong to. Noam Brown is a well known researcher in the field and I see no reason to doubt these claims other than a vague distrust of OpenAI or big tech employees generally which I reject.

Re: OpenAI claims gold-medal performance at IMO 2025

#145

The cynicism/denial on HN about AI is exhausting. Half the comments are some weird form of explaining away the ever increasing performance of these models I've been reading this website for probably 15 years, its never been this bad. many threads are completely unreadable, all the actual educated takes are on X, its almost like there was a talent drain

Probably because both sides have strong vested interests and it’s next to impossible to find a dispassionate point of view. The Pro AI crowd, VC, tech CEOs etc have strong incentive to claim humans are obsolete. Many tech employees see threats to their jobs and want to poopoo any way AI could be useful or competitive.

Or some can spot a euphoric bubble when they see it with lots of participants who have over-invested in 90% of these so called AI startups that are not frontier labs.

Re: OpenAI claims gold-medal performance at IMO 2025

#146
99.99+% of all problems humans face do not require particularly original solutions. Determining whether LLMs can solve truly original (or at least obscure) problems is interesting, and a problem worth solving, but ignores the vast majority of the (near-term at least) impact they will have.

Re: OpenAI claims gold-medal performance at IMO 2025

#147

Earlier quoted context omitted.

While I usually enjoy seeing these discussions, I think they are really pushing the usefulness of bayesian statistics. If one dude says the chance for an outcome is 8% and another says it's 16% and the outcome does occur, they were both pretty wrong, even though it might seem like the one who guessed a few % higher might have had a better belief system. Now if one of them had said 90% while the other said 8% or 16%,…

From a mathematical point of view there are two factors: (1) Initial prior capability of prediction from the human agents and (2) Acceleration in the predicted event. Now we examine the result under such a model and conclude that: The more prior predictive power of human agents imply the more a posterior acceleration of progress in LLMs (math capability). Here we are supposing that the increase in training data is no…

Another take at a sound interpretation:

(1) Bad prior prediction capability of humans imply that result does not provide any information

(2) Good prior prediction capability of humans imply that there is acceleration in math capabilities of LLMs.

Re: OpenAI claims gold-medal performance at IMO 2025

#148
post #76

This is an awesome progress in human achievement to get these machines intelligent. And this is also a fast regress and decline on the human wisdom! We are simply greasing the grooves and letting things slide faster and faster and calling it progress. How does this help to make the human and nature integration better? Does this improve climate or make humans adapt better to changing climate? Are the intelligent machi…

Nobody knows the answers to these questions. Relying on AGI solving problems like climate change seems like a risky strategy but on the other hand it’s very plausible that these tools can help in some capacity. So we have to build, study and find out but also consider any opportunity cost of building these tools versus others.

Solving climate change isn't a technical problem, but a human one. We know the steps we have to take, and have for many years. The hard part is getting people to actually do them.

No human has any idea how to accomplish that. If a machine could, we would all have much to learn from it.

Re: OpenAI claims gold-medal performance at IMO 2025

#149

Earlier quoted context omitted.

- AI competing is "wholly unfair" - "[AI is] far away from being substantially being better than MCTs" ^ pick only one

Yeah it’s a completely fair playing field, it’s completely obvious that AI should be able to compete with humans in the same way that robotics and computers can compete with humanity (and are better suited for many tasks). Whether or not they’re far away from being better than humans is up to debate, but the entire point of these types of benchmarks it to compare them to humans.

>>Yeah it’s a completely fair playing field, it’s completely obvious that AI should be able to compete with humans in the same way that robotics and computers can compete with humanity (and are better suited for many tasks).

Yeah same way computers and robots should be able to win World Chess Championship, 100m dash and Wimbledon.

>>but the entire point of these types of benchmarks it to compare them to humans

The entire point of the competition is to fight against participants who are similar to you, have similar capabilities and go through similar struggles. If you want bot vs human competitions - great - organize it yourself instead of hijacking well established competitions out there.

Re: OpenAI claims gold-medal performance at IMO 2025

#150

The cynicism/denial on HN about AI is exhausting. Half the comments are some weird form of explaining away the ever increasing performance of these models I've been reading this website for probably 15 years, its never been this bad. many threads are completely unreadable, all the actual educated takes are on X, its almost like there was a talent drain

cynacism -> cynicism
Post reply on HN