> level performance on the world’s most prestigious math competition I don't know which one i would consider the most prestigious math competition but it wouldn't be The IMO. The Putnam ranks higher to me and I'm not even an American. But I've come to realise one thing and that is that high-school is very important to Americans...
The Putnam and IMO are quite different. I would suggest the IMO is probably harder...
OpenAI claims gold-medal performance at IMO 2025
551–560 of 737 posts
Re: OpenAI claims gold-medal performance at IMO 2025
#552Earlier quoted context omitted.
The scope and creativity required for IMO is much bigger than chess/GO. Also IMO is taken VERY seriously. It's a huge deal, much bigger than any chess or go tournaments.
Imo competitive math (or programming) is about knowing some tricks and then trying to find a combination of them that works for a given task. The number of tricks and depth required is much less than in go or chess. I don't think it's very creative endeavor in comparison to chess/go. The searching required is less as well. There is a challenge processing natural language and producing solutions in it though. Creativi…
There is no list of tricks that will get a silver much less a gold medal at the IMO. The problem setters try very hard to choose problems that are not just variations of other contests or solvable by routine calculation (indeed some types of problems, like polynomial inequalities, fell out of favor as near-universal techniques made them too routine to well prepared students). Of course there are common themes and patterns that recur--no way around it given the limited curriculum they draw on--but overall I think the IMO does a commendable job at encouraging out-of-the-box thinking within a limited domain. (I've heard a contestant say that IMO prep was memorizing a lot of template solutions, but he was such a genius among geniuses that I think his opinion is irrelevant to the rest of humanity!)
Of course there is always a debate whether competition math reflects skill in research math and other research domains. There's obvious areas of overlap and obvious areas of differences, so it's hard to extrapolate from AI math benchmarks to other domains. But i think it's fair to say the skills needed for the IMO include quite general quantitative reasoning ability, which is very exciting to see LLMs develop.
Re: OpenAI claims gold-medal performance at IMO 2025
#553OpenAI simply can’t be trusted on any benchmarks: https://news.ycombinator.com/item?id=42761648
The proofs are correct, and it's very unlikely that IMO problems were leaked ahead of time. So the options for cheating in this circumstance are that a) IMO are colluding with a few researchers at OpenAI for some reason, or b) @alexwei_ solved the problems himself - both seem pretty unlikely to me.
[1] https://imo2025.au/wp-content/uploads/2025/07/IMO-2025_Closi...
Re: OpenAI claims gold-medal performance at IMO 2025
#554I don't know how much novelty should you expect from IMO every year but i expect many of them be variation of the same problem. These models are trained on all old problem and their various solutions.For LLM models, solving thses problems are as impressive as writing code. There is no high generalization.
Re: OpenAI claims gold-medal performance at IMO 2025
#555Earlier quoted context omitted.
I thought single digit means single significant digit, aka rounding to 10%?
Wasn't 16% the example they were talking about? Isn't that two significant digits? And 16% very much feels ridiculous to a reader when they could've just said 15%.
For what it's worth, I don't think there's anything even slightly wrong with using whatever estimate feels good to you, even if it happens not to fit someone else's criterion for being a nice round number, even if your way of getting the estimate was sticking a finger in the air and saying the first number you thought of. You never make anything more accurate by rounding it[1], and while it's important to keep track of how precise your estimates are I think it's a mistake to try to do that by modifying the numbers. If you have two pieces of information (your best estimate, and how fuzzy it is), you should represent it as two pieces of information[2].
[1] This isn't strictly true, but it's near enough.
[2] Cf. "Pitman's two-bit rule".
Re: OpenAI claims gold-medal performance at IMO 2025
#556Earlier quoted context omitted.
I mean, solutions for the 2025 IMO problems are already available on the internet. How can we be sure these are “unencountered” problems?
They probably have an archived data set from before then that they trained on.
Re: OpenAI claims gold-medal performance at IMO 2025
#557Pre-registering a prediction: When (not if) AI does make a major scientific discovery, we'll hear "well it's not really thinking, it just processed all human knowledge and found patterns we missed - that's basically cheating!"
Re: OpenAI claims gold-medal performance at IMO 2025
#558Earlier quoted context omitted.
What about o3-pro? Remember that their model names make no sense. Edit due to rate-limiting: o3-pro returned an answer after 24 minutes: https://chatgpt.com/share/687bf8bf-c1b0-800b-b316-ca7dd9b009... Whether the CoT amounts to valid mathematical reasoning, I couldn't say, especially because OpenAI models tend to be very cagey with their CoT. Gemini 2.5 Pro seems to have used more sophisticated reasoning ( https://g.…
Only have basic o3 to try. Spent like 10 minutes but did not return any response due to a network error. Checking the thoughts, the model was doing a lot of brute forcing up to n=8, and found k=0,1,3, but no mathematical reasoning was seen.
It convincingly argues that Gemini's answer was wrong, and Gemini agrees ( https://g.co/gemini/share/aa26fb1a4344 ).
So that's pretty cool, IMO. Pitting these two models against each other in a cage match is an underused hack in my experience.
Another observation worth making is that (looking at the Github link) OpenAI didn't just paste an image of the question into the prompt, hit the button and walk away, like I did. They rewrote the prompts carefully to get the best results, and I'm a little surprised people aren't crying foul about that. So I'm pretty impressed with o3-pro's unassisted performance.
Re: OpenAI claims gold-medal performance at IMO 2025
#559I encourage anyone who thinks these are easy high-school problems to try to solve some. They're published (including this year's) at https://www.imo-official.org/problems.aspx . They make my head spin.
[flagged]
Re: OpenAI claims gold-medal performance at IMO 2025
#560Earlier quoted context omitted.
I’m in the US but this was a while back, in the south. It was a highly ranked school and ended up producing lots of PhDs, but many of the families were blue collar and so there just wasn’t any awareness of things like this.
I'm curious on your response to GP's question. Have you heard of AHSME, AMC, or AIME? Nobody mentioned them in high school (1997) until I heard of them online and got my school to participate. 30 kids took the AHSME. Only one qualified for the AIME. And nobody qualified for IMO (though I tell myself I was close). I believe the 1 in a million number.