Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

361–370 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#361
post #160

Earlier quoted context omitted.

I feel you're misunderstanding something. That's not "this exact experiment". Matharena is testing publicly available models against the IMO problem set. OpenAI was announcing the results of a new, unpublished model, on that problems set. It is totally fair to discount OpenAI's statement until we have way more details about their setup, and maybe even until there is some level of public access to the model. But you'r…

Implying results are fraudulent is completely fair when it is a fraud. The previous time they had claims about solving all of the math right there and right then, they were caught owning the company that makes that independent test , and could neither admit nor deny training on closed test set.

Just to quickly clarify:

- OpenAI doesn't own Epoch AI (though they did commission Epoch to make the eval)

- OpenAI denied training on the test set (and further denied training on FrontierMath-derived data, training on data targeting FrontierMath specifically, or using the eval to pick a model checkpoint; in fact, they only downloaded the FrontierMath data after their o3 training set was frozen and they didn't look at o3's FrontierMath results until after the final o3 model was already selected. primary source: https://x.com/__nmca__/status/1882563755806281986)

You can of course accuse OpenAI of lying or being fraudulent, and if that's how you feel there's probably not much I can say to change your mind. One piece of evidence against this is that the primary source linked above no longer works at OpenAI, and hasn't chosen to blow the whistle on the supposed fraud. I work at OpenAI myself, training reasoning models and running evals, and I can vouch that I have no knowledge or hints of any cheating; if I did, I'd probably quit on the spot and absolutely wouldn't be writing this comment.

Totally fine not to take every company's word at face value, but imo this would be a weird conspiracy for OpenAI, with very high costs on reputation and morale.

Re: OpenAI claims gold-medal performance at IMO 2025

#362

The cynicism/denial on HN about AI is exhausting. Half the comments are some weird form of explaining away the ever increasing performance of these models I've been reading this website for probably 15 years, its never been this bad. many threads are completely unreadable, all the actual educated takes are on X, its almost like there was a talent drain

Meh. Some over hype, some under hype. People like you whine and then don't want to listen to any technical concerns. Some of us are implementing things in relation to AI so we know it's not about "increasing performance of models" but actual about the right solution for the right problem. If you think Twitter has "educated takes" then maybe go there and stop being pretentious schmuck over here. Talent drain, lol. I'd…

Both sides are not equally wrong, clearly. Until yesterday prediction markets were saying the probability of an AI getting a gold medal in IMO in 2025 was <20%. So clearly we should be more hyped, not less.

Re: OpenAI claims gold-medal performance at IMO 2025

#363

Noam Brown: > this isn’t an IMO-specific model. It’s a reasoning LLM that incorporates new experimental general-purpose techniques. > it’s also more efficient [than o1 or o3] with its thinking. And there’s a lot of room to push the test-time compute and efficiency further. > As fast as recent AI progress has been, I fully expect the trend to continue. Importantly, I think we’re close to AI substantially contributing…

What's the clear path to improved efficiency now that we've reached peak data?

The thing is, people claimed already a year or two ago that we'd reached peak data and progress would stall since there was no more high-quality human-written text available. Turns out they were wrong, and if anything progress accelerated.

The progress has come from all kinds of things. Better distillation of huge models to small ones. Tool use. Synthetic data (which is not leading to model collapse like theorized). Reinforcement learning.

I don't know exactly where the progress over the next year will be coming from, but it seems hard to believe that we'll just suddenly hit a wall on all of these methods at the same time and discover no new techniques. If progress had slowed down over the last year the wall being near would be a reasonable hypothesis, but it hasn't.

Re: OpenAI claims gold-medal performance at IMO 2025

#364

OpenAI simply can’t be trusted on any benchmarks: https://news.ycombinator.com/item?id=42761648

Somewhat related, but I’ve been feeling as of late what can best be described as “benchmark fatigue”. The latest models can score something like 70% on SWE-bench verified and yet it’s difficult to say what tangible impact this has on actual software development. Likewise, they absolutely crush humans at sport programming but are unreliable software engineers on their own. What does it really mean that an LLM got gold…

[deleted]

Re: OpenAI claims gold-medal performance at IMO 2025

#365
post #9

Some previous predictions: In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025. He thought there was an 8% chance of this happening. Eliezer Yudkowsky said "at least 16%". Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...

Off topic, but am I the only one getting triggered every time I see a rationalist quantify their prediction of the future with single digit accuracy? It's like their magic way of trying to get everyone to forget that they reached their conclusion in completely hand-wavy way, just like every other human being. But instead of saying "low confidence" or "high confidence" like the rest of us normies, they will tell you t…

No, you are right, this hyper-numericalism is just astrology for nerds.

Re: OpenAI claims gold-medal performance at IMO 2025

#366
post #353

Earlier quoted context omitted.

The correctness of 8%, 16%, and 90% are all equally unknown since we only have one timeline, no?

That's why you have to let these people make predictions about many things. Than you can weigh the 8, 16, and 90 pct and see who is talking out of their ass.

That's just the frequentist approach. But we're talking about bayesian statistics here.

Re: OpenAI claims gold-medal performance at IMO 2025

#367

Earlier quoted context omitted.

[flagged]

I am a professor in a math department (I teach statistics but there is a good complement of actual math PhDs) and there are only about 10% who care about these types of problems and definitely less than half who could get gold on an IMO test even if they didn’t care. They are all outstanding mathematicians, but the IMO type questions are not something that mathematicians can universally solve without preparation. The…

> They are all outstanding mathematicians, but the IMO type questions are not something that mathematicians can universally solve without preparation.

So IMO is basically the leetcode of Mathematics.

Re: OpenAI claims gold-medal performance at IMO 2025

#368
post #9

Some previous predictions: In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025. He thought there was an 8% chance of this happening. Eliezer Yudkowsky said "at least 16%". Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...

Off topic, but am I the only one getting triggered every time I see a rationalist quantify their prediction of the future with single digit accuracy? It's like their magic way of trying to get everyone to forget that they reached their conclusion in completely hand-wavy way, just like every other human being. But instead of saying "low confidence" or "high confidence" like the rest of us normies, they will tell you t…

There is no honor in hiding behind euphemisms. Rationalists say ‘low confidence’ and ‘high confidence’ all the time, just not when they're making an actual bet and need to directly compare credences. And the 16.27% mockery is completely dishonest. They used less than a single significant figure.

Re: OpenAI claims gold-medal performance at IMO 2025

#369
post #345

Earlier quoted context omitted.

> I think we’re close to AI substantially contributing to scientific discovery. The new "Full Self-Driving next year"?

"AI" already contributes "substantially" to "scientific discovery". It's a very safe statement to make, whereas "full self-driving" has some concrete implications.

Well I also think full self-driving contribute substantially to navigating the car on the street..

Re: OpenAI claims gold-medal performance at IMO 2025

#370

Earlier quoted context omitted.

How is a claim , "clear evidence" to anything?

Most evidence you have about the world is claims from other people, not direct experiment. There seems to be a thought-terminating cliche here on HN, dismissing any claim from employees of large tech companies. Unlike seemingly most here on HN, I judge people's trustworthiness individually and not solely by the organization they belong to. Noam Brown is a well known researcher in the field and I see no reason to doub…

A thought-terminating cliché? Not at all, certainly not when it comes to claims of technological or scientific breakthroughs. After all, that's partly why we have peer review and an emphasis on reproducibility. Until such a claim has been scrutinised by experts or reproduced by the community at large, it remains an unverified claim.

>> Unlike seemingly most here on HN, I judge people's trustworthiness individually and not solely by the organization they belong to.

That has nothing to do with anything I said. A claim can be false without it being fraudulent, in fact most false claims are probably not fraudulent; though, still, false.

Claims are also very often contested. See e.g. the various claims of Quantum Superiority and the debate they have generated.

Science is a debate. If we believe everything anyone says automatically, then there is no debate.

Post reply on HN