Terence Tao on the matter - https://imgur.com/a/terence-tao-on-supposed-gold-imo-sMKP0bm
OpenAI claims gold-medal performance at IMO 2025
661–670 of 737 posts
Re: OpenAI claims gold-medal performance at IMO 2025
#662Earlier quoted context omitted.
It is not "supremacist" to believe that depriving hundreds of millions of people from higher education in their native language is deeply unjust. This reflection was prompted by a comment on why Indian languages are not represented in international competitions, which was prompted by a comment on the competition being available in many languages. Discussions online have a tendency to go off into tangents like this. I…
Much more efficient for us to all speak the same language. Trying to create fragmentation is inefficient.
Re: OpenAI claims gold-medal performance at IMO 2025
#663Earlier quoted context omitted.
This is slightly tedious to do by hand but there isn't really anything interesting going on in that problem - it's just solving a quadratic equation over the complex numbers.
That isn't much of an argument; nothing in math is truly interesting if you take that approach. exp(i\pi)+1=0 could be said to be dis-interesting because it is just rotation on the complex plane. But it is the opposite - it is interesting because it turned out to be rotation on the complex plane but approached from summing infinite series. Similarly you can say that solving a quadratic over complex numbers is dis-int…
This is distinct both from other typical IMO problems that I've seen and from research mathematics which usually do require some amount of creativity.
> exp(i\pi)+1=0
If your definition of "exp(i*theta)" is literally "rotation of the number 1 by theta degrees counterclockwise", then indeed what you quoted is a triviality and contains no nugget of insight (how could it?).
It becomes nontrivial when your definition of "exp" is any of the following:
- The everywhere absolutely convergent power series sum_{i=0}^\infty z^n/n!
- The unique function solving the IVP y'=y, y(0)=1
- The unique holomorphic extension of the real-valued exponential function to the complex numbers
Going from any of these definitions to "exp(i*\pi)+1=0" from scratch requires quite a bit of clever mathematics (such as proving convergence of the various series, comparing terms, deriving the values of sin and cos at pi from their power series representation, etc.). That's definitely not something that a motivated high schooler would be able to derive from scratch.
Re: OpenAI claims gold-medal performance at IMO 2025
#664Earlier quoted context omitted.
All the past IMO problems are known to public and contestants practice on them. If solving an IMO problem is the simple matter of "looking at all the past problems and apply the same pattern," you'd expect human contestants to do a lot better.
I think you haven't gone thru AMC8, AMC10, AIME competitions. If you are so confident, try giving an unsolved math problem outside the high school math competitions.
Re: OpenAI claims gold-medal performance at IMO 2025
#665> level performance on the world’s most prestigious math competition I don't know which one i would consider the most prestigious math competition but it wouldn't be The IMO. The Putnam ranks higher to me and I'm not even an American. But I've come to realise one thing and that is that high-school is very important to Americans...
Re: OpenAI claims gold-medal performance at IMO 2025
#666In the RLHF sphere you could tell some AI company/companies were targeting this because of how many IMO RLHF’ers they were hiring specifically. I don’t think it’s really easy to say how much “progress” this is given that.
Re: OpenAI claims gold-medal performance at IMO 2025
#667From that thread: "The model solved P1 through P5; it did not produce a solution for P6." It's interesting that it didn't solve the problem that was by far the hardest for humans too. China, the #1 team got only 21/42 points on it. In most other teams nobody solved it.
In the IMO, the idea is that the first day you get P1, P2 and P3, and the second day you get P4, P5 and P6. Usually, ordered by difficulty, they are P1, P4, P2, P5, P3, P6. So, usually P1 is "easy" and P6 is very hard. At least that is the intended order, but sometime reality disagree. Edit: Fixed P4 -> P3. Thanks.
Day 1: P1 P3 P5 (odds)
Day 2: P2 P4 P6 (evens)
Then the problem # is the difficulty.
Re: OpenAI claims gold-medal performance at IMO 2025
#668Earlier quoted context omitted.
I doubt this is coming from RLHF - tweets from the lead researcher state that this result flows from a research breakthrough which enables RLVR on less verifiable domains.
Math RLHF already has verifiable ground truth/right vs wrong, so I don't what this distinction really shows. And AI changes so quickly that there is a breakthrough every week. Call my cynical, but I think this is an RLHF/RLVR push in a narrow area--IMO was chosen as a target and they hired specifically to beat this "artificial" target.
Re: OpenAI claims gold-medal performance at IMO 2025
#669Re: OpenAI claims gold-medal performance at IMO 2025
#670Terence Tao on the matter - https://imgur.com/a/terence-tao-on-supposed-gold-imo-sMKP0bm
Actual post instead of ad-decorated screnshot: https://mathstodon.xyz/@tao/114881418225852441 (thread continued in https://mathstodon.xyz/@tao/114881419368778558 and https://mathstodon.xyz/@tao/114881420636881657 ).
It’s as-if we had learned whale song, and then within two years a whale had won a Nobel prize for their research in high pressure aquatic environments. You’d similarly get naysayers debating the finer points of what special advantage whales may have in that particular field, neglecting the stunned shock of the general population — “Whales are publishing research papers now!? Award winning papers at that!?”