Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

11–20 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#11
Definitely interesting. Two thoughts. First, are the IMO questions somewhat related to other openly available questions online, making it easier for LLMs that are more efficient and better at reasoning to deduce the results from the available content?

Second, happy to test it on open math conjectures or by attempting to reprove recent math results.

Re: OpenAI claims gold-medal performance at IMO 2025

#12
post #7

Progress is astounding. Recently report published about evaluation of LLMs on IMO 2025. o3 high didn't even get bronze. https://matharena.ai/imo/ Waiting for Terry Tao's thoughts, but these kind of things are good use of AI. We need to make science progress faster rather than disrupting our economy without being ready.

[flagged]

I mean progress speed, few months ago they released o3 it has 16 pt in imo 2025

Re: OpenAI claims gold-medal performance at IMO 2025

#13
post #11

Definitely interesting. Two thoughts. First, are the IMO questions somewhat related to other openly available questions online, making it easier for LLMs that are more efficient and better at reasoning to deduce the results from the available content? Second, happy to test it on open math conjectures or by attempting to reprove recent math results.

You mean as in the previous years questions will have been used to train it? Yes, they are the same questions and due to them limited format on math questions, there are repeats so LLMs should fundamentally be able to recognise a structure and similarities and use that.

Re: OpenAI claims gold-medal performance at IMO 2025

#15
post #11

Definitely interesting. Two thoughts. First, are the IMO questions somewhat related to other openly available questions online, making it easier for LLMs that are more efficient and better at reasoning to deduce the results from the available content? Second, happy to test it on open math conjectures or by attempting to reprove recent math results.

From what I've seen, IMO question sets are very diverse. Moreover, humans also train on all available set of math olympiad questions and similar sets too. It seems fair game to have the AI train on them as well.

For 2, there's an army of independent mathematicians right now using automated theorem provers to formalise more or less all mathematics as we know it. It seems like open conjectures are chiefly bounded by a genuine lack of new tools/mathematics.

Re: OpenAI claims gold-medal performance at IMO 2025

#16
These are high school level only in the sense of assumed background knowledge, they are extremely difficult.

Professional mathematicians would not get this level of performance, unless they have a background in IMO themselves.

This doesn’t mean that the model is better than them in math, just that mathematicians specialize in extending the frontier of math.

The answers are not in the training data.

This is not a model specialized to IMO problems.

Re: OpenAI claims gold-medal performance at IMO 2025

#17
post #3

Earlier quoted context omitted.

[flagged]

Which would be impressive if we knew those problems weren't in the training data already. I mean it is quite impressive how language models are able to mobilize the knowledge they have been trained on, especially since they are able to retrieve information from sources that may be formatted very differently, with completely different problem statement sentences, different variable names and so on, and really operate…

Considering these LLM utilise the entirety of the internet, there will be no unique problems that come up in the oLympiad. Even across the course of a degree, you will have likely been exposed to 95% of the various ways to write problems. As you say, retrieval is really the only skill here. There is likely no reasoning.

Re: OpenAI claims gold-medal performance at IMO 2025

#19

These are high school level only in the sense of assumed background knowledge, they are extremely difficult. Professional mathematicians would not get this level of performance, unless they have a background in IMO themselves. This doesn’t mean that the model is better than them in math, just that mathematicians specialize in extending the frontier of math. The answers are not in the training data. This is not a mode…

[flagged]

Re: OpenAI claims gold-medal performance at IMO 2025

#20
post #18
post #8

Earlier quoted context omitted.

[flagged]

Velocity of AI progress in recent years is exceeded only by velocity of goalposts.

The goalposts should focus on being able to make a coherent statement using papers on a subject with sources. At this point it can't do that for any remotely cutting edge topic. This is just a distraction.
Post reply on HN