OpenAI claims gold-medal performance at IMO 2025
671–680 of 737 posts
Re: OpenAI claims gold-medal performance at IMO 2025
#672Earlier quoted context omitted.
This is slightly tedious to do by hand but there isn't really anything interesting going on in that problem - it's just solving a quadratic equation over the complex numbers.
That isn't much of an argument; nothing in math is truly interesting if you take that approach. exp(i\pi)+1=0 could be said to be dis-interesting because it is just rotation on the complex plane. But it is the opposite - it is interesting because it turned out to be rotation on the complex plane but approached from summing infinite series. Similarly you can say that solving a quadratic over complex numbers is dis-int…
If you go to the complex plane, you are re-defining the plane. If you redefine the plane, then you can do anything. The puzzle is about confusing the observer who is expecting a solution in a certain dimension.
Re: OpenAI claims gold-medal performance at IMO 2025
#673Earlier quoted context omitted.
To me, this is a tell of human-involvement in the model solution. There is no reason why machines would do badly on exactly the problem which humans do badly as well - without humans prodding the machine towards a solution. Also, there is no reason why machines could not produce a partial or wrong answer to problem 6 which seems like survivor bias to me. ie, that only correct solutions were cherrypicked.
Maybe it’s a hint that our current training techniques can create models comparable to the best humans in a given subject, but that’s the limit.
IMO is not the best humans in a given subject, college competitions are at a much higher level than high school competitions, and you have even higher level above that since college competitions are still limited to students.
Re: OpenAI claims gold-medal performance at IMO 2025
#674Earlier quoted context omitted.
They also said this is not part of GPT-5, and “will be released later”. It’s very, very likely a model specifically fine-tuned for this benchmark, where afterwards they’ll evaluate what actual real-world problems it’s good at (eg like “use o4-mini-high for coding”).
Humans who excel at IMO questions are also "fine tuned" on them in the sense that they practice them for hundreds of hours
So its a big difference if you use a general intelligence system and makes it do well in math, or when you create a specialized system that is only good at math and can't be used to get good in other areas.
Re: OpenAI claims gold-medal performance at IMO 2025
#675Google also joined IMO, and got gold prize. https://x.com/natolambert/status/1946569475396120653 OAI announced early, probably we will hear announcement from Google soon.
Google’s AlphaProof, which got a silver last year, has been using a neural symbolic approach. This gold from OpenAI was pure LLM. We’ll have to see what Google announces, but the LLM approach is interesting because it will likely generalize to all kinds of reasoning problems, not just mathematical proofs.
Re: OpenAI claims gold-medal performance at IMO 2025
#676Earlier quoted context omitted.
Actual post instead of ad-decorated screnshot: https://mathstodon.xyz/@tao/114881418225852441 (thread continued in https://mathstodon.xyz/@tao/114881419368778558 and https://mathstodon.xyz/@tao/114881420636881657 ).
Fair points, but the reason everyone is amazed is that five years ago this was entirely impossible for computers irrespective of the competition format or rules. It’s as-if we had learned whale song, and then within two years a whale had won a Nobel prize for their research in high pressure aquatic environments. You’d similarly get naysayers debating the finer points of what special advantage whales may have in that…
A computer system that can perform these tasks that were unthinkably complex a few years ago is quite impressive. That is a big win, and it can be celebrated. They don’t need to be celebrated as a “gold medalist” if they didn’t perform according to the same criteria as a gold-medalist.
Re: OpenAI claims gold-medal performance at IMO 2025
#677Earlier quoted context omitted.
> I actually think this “cheating” is fine. In fact it’s preferable. The thing with IMO, is the solutions are already known by someone . So suppose the model got the solutions beforehand, and fed them into the training model. Would that be an acceptable level of "cheating" in your view?
Surely you jest. The cheating would be the same cheating as any other situation - someone inside the IMO skipping the questions and answers to people outside then that being used to compete. Fine - but why? If this were discovered then it would be disastrous for everyone involved, and for what? A noteworthy HN link? The downside would be international scandal and careers destroyed. The upside is imperceptible. Finall…
Re: OpenAI claims gold-medal performance at IMO 2025
#678Earlier quoted context omitted.
The book Superforecasting documented that for their best forecasters, rounding off that last percent would reliably reduce Brier scores. Whether rationalists who are publicly commenting actually achieve that level of reliability is an open question. But that humans can be reliable enough in the real world that the last percentage matters, has been demonstrated.
Your comment is incredibly confusing (possibly misleading) because of the key details you've omitted. > The book Superforecasting documented that for their best forecasters, rounding off that last percent would reliably reduce Brier scores. Rounding off that last percent... to what , exactly? Are you excluding the exceptions I mentioned (i.e. when you're already close to 0% or 100%?) Nobody is arguing that 3% -> 4% i…
Re: OpenAI claims gold-medal performance at IMO 2025
#679Earlier quoted context omitted.
Your comment is incredibly confusing (possibly misleading) because of the key details you've omitted. > The book Superforecasting documented that for their best forecasters, rounding off that last percent would reliably reduce Brier scores. Rounding off that last percent... to what , exactly? Are you excluding the exceptions I mentioned (i.e. when you're already close to 0% or 100%?) Nobody is arguing that 3% -> 4% i…
To the nearest 5%, for percentages in that middle range. It is not just 16% -> 15%. But also 46% -> 45%.