Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

371–380 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#371
post #9

Some previous predictions: In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025. He thought there was an 8% chance of this happening. Eliezer Yudkowsky said "at least 16%". Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...

Off topic, but am I the only one getting triggered every time I see a rationalist quantify their prediction of the future with single digit accuracy? It's like their magic way of trying to get everyone to forget that they reached their conclusion in completely hand-wavy way, just like every other human being. But instead of saying "low confidence" or "high confidence" like the rest of us normies, they will tell you t…

Interestingly, this is actually a question that's been looked at empirically!

Take a look at this paper: https://scholar.harvard.edu/files/rzeckhauser/files/value_of...

They took high-precision forecasts from a forecasting tournament and rounded them to coarser buckets (nearest 5%, nearest 10%, nearest 33%), to see if the precision was actually conveying any real information. What they found is that if you rounded the forecasts of expert forecasters, Brier scores got consistently worse, suggesting that expert forecast precision at the 5% level is still conveying useful, if noisy, information. They also found that less expert forecasters took less of a hit from rounding their forecasts, which makes sense.

It's a really interesting paper, and they recommend that foreign policy analysts try to increase precision rather than retreating to lumpy buckets like "likely" or "unlikely".

Based on this, it seems totally reasonable for a rationalist to make guesses with single digit precision, and I don't think it's really worth criticizing.

Re: OpenAI claims gold-medal performance at IMO 2025

#372
post #363

Earlier quoted context omitted.

What's the clear path to improved efficiency now that we've reached peak data?

The thing is, people claimed already a year or two ago that we'd reached peak data and progress would stall since there was no more high-quality human-written text available. Turns out they were wrong, and if anything progress accelerated. The progress has come from all kinds of things. Better distillation of huge models to small ones. Tool use. Synthetic data (which is not leading to model collapse like theorized).…

I'm loving it, can't wait to deploy this stuff locally. The mainframe will be replaced by commodity hardware, OpenAI will stare down the path of IBM unless they reinvent themselves.

Re: OpenAI claims gold-medal performance at IMO 2025

#373
post #186

Earlier quoted context omitted.

Latent?? If you looked at RLHF hiring over the last year, there was a huge hiring of IMO competitors to RLHF. This was a new, highly targeted, highly funded RLHF’ing.

Can you provide any kind of source? Very curious about this!

https://work.mercor.com/jobs/list_AAABljpKHPMmFMXrg2VM0qz4

https://benture.io/job/international-math-olympiad-participa...

https://job-boards.greenhouse.io/xai/jobs/4538773007

And Outlier/Scale, which was bought by Meta (via Scale), had many IMO-required Math AI trainer jobs on LinkedIn. I can't find those historical ones though.

I'm just one piece in the cog and this is an anecdote, but there was a huge upswing in IMO or similar RLHF job postings over the past 6mo-year.

Re: OpenAI claims gold-medal performance at IMO 2025

#374

Noam Brown: > this isn’t an IMO-specific model. It’s a reasoning LLM that incorporates new experimental general-purpose techniques. > it’s also more efficient [than o1 or o3] with its thinking. And there’s a lot of room to push the test-time compute and efficiency further. > As fast as recent AI progress has been, I fully expect the trend to continue. Importantly, I think we’re close to AI substantially contributing…

That's a big leap from "answering test questions" to "contributing to scientific discovery".

Re: OpenAI claims gold-medal performance at IMO 2025

#375

Earlier quoted context omitted.

It almost certainly is specialized to IMO problems, look at the way it is answering the questions: https://xcancel.com/alexwei_/status/1946477742855532918 E.g here: https://pbs.twimg.com/media/GwLtrPeWIAUMDYI.png?name=orig Frankly it looks to me like it's using an AlphaProof style system, going between natural language and Lean/etc. Of course OpenAI will not tell us any of this.

I actually think this “cheating” is fine. In fact it’s preferable. I don’t need an AI that can act as a really expensive calculator or solver. We’ve already built really good calculators and solvers that are near optimal. What has been missing is the abductive ability to successfully use those tools in an unconstrained space with agency. I find really no value in avoiding the optimal or near optimal techniques we’ve…

> I actually think this “cheating” is fine. In fact it’s preferable.

The thing with IMO, is the solutions are already known by someone.

So suppose the model got the solutions beforehand, and fed them into the training model. Would that be an acceptable level of "cheating" in your view?

Re: OpenAI claims gold-medal performance at IMO 2025

#376

Earlier quoted context omitted.

> We can only go off their word We’re talking about Sam Altman’s company here. The same company that started out as a non profit claiming they wanted to better the world. Suggesting they should be given the benefit of the doubt is dishonest at this point.

“they must be lying because I personally dislike them” This is why HN threads about AI have become exhausting to read

Yeah, that's how the concept of "reputation" works.

Re: OpenAI claims gold-medal performance at IMO 2025

#377

Noam Brown: > this isn’t an IMO-specific model. It’s a reasoning LLM that incorporates new experimental general-purpose techniques. > it’s also more efficient [than o1 or o3] with its thinking. And there’s a lot of room to push the test-time compute and efficiency further. > As fast as recent AI progress has been, I fully expect the trend to continue. Importantly, I think we’re close to AI substantially contributing…

What's the clear path to improved efficiency now that we've reached peak data?

there is also huge realm of private/commercial data which is not absorbed by LLMs yet. I think there are way more private/commercial data than public data.

Re: OpenAI claims gold-medal performance at IMO 2025

#378

Earlier quoted context omitted.

While I usually enjoy seeing these discussions, I think they are really pushing the usefulness of bayesian statistics. If one dude says the chance for an outcome is 8% and another says it's 16% and the outcome does occur, they were both pretty wrong, even though it might seem like the one who guessed a few % higher might have had a better belief system. Now if one of them had said 90% while the other said 8% or 16%,…

The correctness of 8%, 16%, and 90% are all equally unknown since we only have one timeline, no?

If one is calibrated to report proper percentages and assigns 8% to 25 distinct events, you should expect 2 of the events to occur; 4 in case of 16% and 22.5 in case of 90%. Assuming independence (as is sadly too often done) standard math of binomial distributions can be applied and used to distinguish the prediction's accuracy probabilistically despite no actual branching or experimental repetition taking place.

Re: OpenAI claims gold-medal performance at IMO 2025

#379

Interesting that the proofs seem to use a limited vocabulary: https://github.com/aw31/openai-imo-2025-proofs/blob/main/pro... Why waste time say lot word when few word do trick :) Also worth pointing out that Alex Wei is himself a gold medalist at IOI.

Are you saying "see the world?" or "seaworld"?

Re: OpenAI claims gold-medal performance at IMO 2025

#380

I am quite surprised that Deepmind with MCTS wasnt able to figure out math performance itself.

Google will also have good results to report for this year's IMO, OpenAI just beat them to the announcement

I think google did some official collaboration with IMO, and will announce later. Or at least that's what I read from the IMO official saying "AI companies should wait 1 week before announcing so that we can celebrate the human winners" and "to my knowledge oai was not officially collaborating with IMO" ...
Post reply on HN