Some previous predictions: In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025. He thought there was an 8% chance of this happening. Eliezer Yudkowsky said "at least 16%". Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...
OpenAI claims gold-medal performance at IMO 2025
91–100 of 737 posts
Re: OpenAI claims gold-medal performance at IMO 2025
#92Re: OpenAI claims gold-medal performance at IMO 2025
#93[flagged]
Re: OpenAI claims gold-medal performance at IMO 2025
#94Definitely interesting. Two thoughts. First, are the IMO questions somewhat related to other openly available questions online, making it easier for LLMs that are more efficient and better at reasoning to deduce the results from the available content? Second, happy to test it on open math conjectures or by attempting to reprove recent math results.
You mean as in the previous years questions will have been used to train it? Yes, they are the same questions and due to them limited format on math questions, there are repeats so LLMs should fundamentally be able to recognise a structure and similarities and use that.
Re: OpenAI claims gold-medal performance at IMO 2025
#95Wow. That's an impressive result, but how did they do it? Wei references scaling up test-time compute, so I have to assume they threw a boatload of money at this. I've heard talk of running models in parallel and comparing results - if OpenAI ran this 10000 times in parallel and cherry-picked the best one, this is a lot less exciting. If this is legit, then we need to know what tools were used and how the model used…
You don’t have to be hyped to be amazed. You can retain the ability to dream while not buying into the snake oil. This is amazing no matter what ensemble of techniques used. In fact - you should be excited if we’ve started to break out of the limitations of forcing NN to be load bearing in literally everything. That’s a sign of maturing technology not of limitations.
Re: OpenAI claims gold-medal performance at IMO 2025
#96OpenAI simply can’t be trusted on any benchmarks: https://news.ycombinator.com/item?id=42761648
Re: OpenAI claims gold-medal performance at IMO 2025
#97Some previous predictions: In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025. He thought there was an 8% chance of this happening. Eliezer Yudkowsky said "at least 16%". Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...
Impressive prediction, especially pre-ChatGPT. Compare to Gary Marcus 3 months ago: https://garymarcus.substack.com/p/reports-of-llms-mastering-... We may certainly hope Eliezer's other predictions don't prove so well-calibrated.
Re: OpenAI claims gold-medal performance at IMO 2025
#98From that thread: "The model solved P1 through P5; it did not produce a solution for P6." It's interesting that it didn't solve the problem that was by far the hardest for humans too. China, the #1 team got only 21/42 points on it. In most other teams nobody solved it.
In the IMO, the idea is that the first day you get P1, P2 and P3, and the second day you get P4, P5 and P6. Usually, ordered by difficulty, they are P1, P4, P2, P5, P3, P6. So, usually P1 is "easy" and P6 is very hard. At least that is the intended order, but sometime reality disagree. Edit: Fixed P4 -> P3. Thanks.
Re: OpenAI claims gold-medal performance at IMO 2025
#99Progress is astounding. Recently report published about evaluation of LLMs on IMO 2025. o3 high didn't even get bronze. https://matharena.ai/imo/ Waiting for Terry Tao's thoughts, but these kind of things are good use of AI. We need to make science progress faster rather than disrupting our economy without being ready.
[flagged]
Every time an LLM reaches a new benchmark there’s a scramble to downplay it and move the goalposts for what should be considered impressive.
The International Math Olympiad was used by many people as an example of something that would be too difficult for LLMs. It has been a topic of discussion for some time. The fact that an LLM has achieved this level of performance is very impressive.
You’re downplaying the difficulty of these problems. It’s called international because the best in the entire world are challenged by it.
Re: OpenAI claims gold-medal performance at IMO 2025
#100Earlier quoted context omitted.
Impressive prediction, especially pre-ChatGPT. Compare to Gary Marcus 3 months ago: https://garymarcus.substack.com/p/reports-of-llms-mastering-... We may certainly hope Eliezer's other predictions don't prove so well-calibrated.
Gary Marcus is so systematically and overconfidently wrong that I wonder why we keep talking about this clown.