Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

51–60 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#52

Wow. That's an impressive result, but how did they do it? Wei references scaling up test-time compute, so I have to assume they threw a boatload of money at this. I've heard talk of running models in parallel and comparing results - if OpenAI ran this 10000 times in parallel and cherry-picked the best one, this is a lot less exciting. If this is legit, then we need to know what tools were used and how the model used…

I don't think it's much less exciting if they ran it 10000 parallel? It implies an ability to discern when the proof is correct and rigorous (which o3 can't do consistently) and also means that outputting the full proof is within capabilities even if rare.

Re: OpenAI claims gold-medal performance at IMO 2025

#53
post #43

Earlier quoted context omitted.

Nope https://x.com/polynoamial/status/1946478249187377206?s=46&t=...

If you don't have a Twitter account then x.com links are useless, use a mirror: https://xcancel.com/polynoamial/status/1946478249187377206 Anyway, that doesn't refute my point, it's just PR from a weaselly and dishonest company. I didn't say it was "IMO-specific" but the output strongly suggests specialized tooling and training, and they said this was an experimental LLM that wouldn't be released. I strongly suspect…

We can only go off their word unfortunately and they say no formal math. so I assume it's being eval'd by a verifier model instead of a formal system. There's actually some hints of this b/c geometry in Lean is not that well developed so unless they also built their own system it's hard to do it formally (though their P2 proof is by coordinate bash (computation by algebra instead of geometric construction) so it's hard to tell.

Re: OpenAI claims gold-medal performance at IMO 2025

#54

Wow. That's an impressive result, but how did they do it? Wei references scaling up test-time compute, so I have to assume they threw a boatload of money at this. I've heard talk of running models in parallel and comparing results - if OpenAI ran this 10000 times in parallel and cherry-picked the best one, this is a lot less exciting. If this is legit, then we need to know what tools were used and how the model used…

> what tools were used and how the model used them

According to the twitter thread, the model was not given access to tools.

Re: OpenAI claims gold-medal performance at IMO 2025

#55

Earlier quoted context omitted.

> high school/early university maths problems should not have been a stretch at all for it This is a ridiculous understatement of the difficulty of getting gold at the IMO.

[flagged]

There are entire fields of math with exceptional people trying to solve impossibly hard problems that utilize quite literally 0 calculus.

Many of them are also questions that eventually end up with proofs or solutions that only require very high level understanding of basic principles. But when I say very high I mean like impossibly high for the average person and ability to combine simple concepts to solve complex problems.

I'd wager the majority of Math graduates from universities would struggle to answer most IMO questions.

Re: OpenAI claims gold-medal performance at IMO 2025

#57

Earlier quoted context omitted.

> high school/early university maths problems should not have been a stretch at all for it This is a ridiculous understatement of the difficulty of getting gold at the IMO.

[flagged]

It's like saying getting a gold medal in boxing is not hard, because it doesn't involve any firearms

Re: OpenAI claims gold-medal performance at IMO 2025

#58

Earlier quoted context omitted.

Getting gold at the IMO is pretty damn hard. I grew up in a relatively underserved rural city. I skipped multiple grades in math, completed the first two years of college math classes while in high school, and won the award for being the best at math out of everyone in my school. I've met and worked with a few IMO gold medalists. Even though I was used to scoring in the 99th percentile on all my tests, it felt like t…

The trouble is, getting an IMO gold medal is much easier (by frequency) than being the #1 Go player in the world, which was achieved by AI 10 years ago. I'm not sure it's enough to just gesture at the task; drilling down into precisely how it was achieved feels important. (Not to take away from the result, which I'm really impressed by!)

The "AI" that won Go was Monte Carlo tree search on a neural net "memory" of the outcome of millions of previous games; this is a LLM solving open ended problems. The tasks are hardly even comparable.

Re: OpenAI claims gold-medal performance at IMO 2025

#59

Earlier quoted context omitted.

The trouble is, getting an IMO gold medal is much easier (by frequency) than being the #1 Go player in the world, which was achieved by AI 10 years ago. I'm not sure it's enough to just gesture at the task; drilling down into precisely how it was achieved feels important. (Not to take away from the result, which I'm really impressed by!)

The "AI" that won Go was Monte Carlo tree search on a neural net "memory" of the outcome of millions of previous games; this is a LLM solving open ended problems. The tasks are hardly even comparable.

And then they created AlphaGo Zero, which is not trained on any previous games, and it was even stronger!

https://deepmind.google/discover/blog/alphago-zero-starting-...

Re: OpenAI claims gold-medal performance at IMO 2025

#60

Earlier quoted context omitted.

> high school/early university maths problems should not have been a stretch at all for it This is a ridiculous understatement of the difficulty of getting gold at the IMO.

[flagged]

Olympiad questions don't require advanced concepts except maybe some classical geometry techniques that you wouldn't normally encounter in modern research mathematics. But they're fundamentally designed as puzzles. You need to spot the tricks.
Post reply on HN