OpenAI claims gold-medal performance at IMO 2025
101–110 of 737 posts
Re: OpenAI claims gold-medal performance at IMO 2025
#102BTW; “Gold medal performance “ looks a promotional term for me.
Re: OpenAI claims gold-medal performance at IMO 2025
#103However, I expect that geometric intuition may still be lacking mostly because of the difficulty of encoding it in a form which an LLM can easily work with. After all, Chatgpt still can't draw a unicorn [1] although it seems to be getting closer.
Re: OpenAI claims gold-medal performance at IMO 2025
#104Not sure there is a good writeup about it yet but here is the livestream: https://www.youtube.com/live/TG3ChQH61vE.
Re: OpenAI claims gold-medal performance at IMO 2025
#105[flagged]
Re: OpenAI claims gold-medal performance at IMO 2025
#106Now it is just doing a bunch of tweets?
Re: OpenAI claims gold-medal performance at IMO 2025
#107In fact no car company claims “gold medal” performance in Olympic running even they can do that 100 yeas ago. Obviously since IMO does not generate much money so it is an easy target. BTW; “Gold medal performance “ looks a promotional term for me.
Re: OpenAI claims gold-medal performance at IMO 2025
#108Wow. That's an impressive result, but how did they do it? Wei references scaling up test-time compute, so I have to assume they threw a boatload of money at this. I've heard talk of running models in parallel and comparing results - if OpenAI ran this 10000 times in parallel and cherry-picked the best one, this is a lot less exciting. If this is legit, then we need to know what tools were used and how the model used…
Why is that less exciting? A machine competing in an unconstrained natural language difficult math contest and coming out on top by any means is breath taking science fiction a few years ago - now it’s not exciting? Regardless of the tools for verification or even solvers - why is the goal post moving so fast? There is no bonus for “purity of essence” and using only neural networks. We live in an era where it’s hard…
Certainly the emergent behaviour is exciting but we tend to jump to conclusions as to what it implies.
This means we are far more trusting with software that lacks formal guarantees than we should be. We are used to software being sound by default but otherwise a moron that requires very precise inputs and parameters and testing to act correctly. System 2 thinking.
Now with NN it's inverted: it's a brilliant know-it-all but it bullshits a lot, and falls apart in ways we may gloss over, even with enormous resources spent on training. It's effectively incredible progress on System 1 thinking with questionable but evolving System 2 skills where we don't know the limits.
If you're not familiar with System 1 / System 2, it's googlable .
Re: OpenAI claims gold-medal performance at IMO 2025
#109[flagged]
- "[AI is] far away from being substantially being better than MCTs"
^ pick only one
Re: OpenAI claims gold-medal performance at IMO 2025
#110Am I missing something or is this completely meaningless? It's 100% opaque, no details whatsoever and no transparency or reproducibility. I wouldn't trust these results as it is. Considering that there are trillions of dollars on the line as a reward for hyping up LLMs, I trust it even less.