Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

451–460 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#451

>AI model performs astounding feat everyone claimed was impossible or won’t be achieved for a while >Commenters on HN claim it must not be that hard, or OpenAI is lying, or cheated. Anything but admit that it is impressive Every time on this site lol. A lot of people here have an emotional aversion to accepting AI progress. They’re deep in the bargaining/anger/denial phase.

I guess my major question would be: does the training data include anything from 2025 which may have included information about the IMO 2025? Given that AI companies are constantly trying to slurp up any and all data online, if the model was derived from existing work, it's maybe less impressive than at first glance. If present-day model does well at IMO 2026, that would be nice.

Are the human participants in the IMO held to the same standard?

Re: OpenAI claims gold-medal performance at IMO 2025

#452

I tried P1 on chatgpt-o4-high, it tells me the solution is k=0 or 1. It doesn’t even know that k=3 is a solution for n=3. Such a solution would get 0/7 in the actual IMO.

What about o3-pro? Remember that their model names make no sense.

Edit due to rate-limiting:

o3-pro returned an answer after 24 minutes: https://chatgpt.com/share/687bf8bf-c1b0-800b-b316-ca7dd9b009... Whether the CoT amounts to valid mathematical reasoning, I couldn't say, especially because OpenAI models tend to be very cagey with their CoT.

Gemini 2.5 Pro seems to have used more sophisticated reasoning ( https://g.co/gemini/share/c325915b5583 ) but it got a slightly different answer. Its chain of thought was unimpressive to say the least, so I'm not sure how it got its act together for the final explanation.

Claude Opus 4 appears to have missed the main solution the others found: https://claude.ai/share/3ba55811-8347-4637-a5f0-fd8790aa820b

Be interesting if someone could try Grok 4.

Re: OpenAI claims gold-medal performance at IMO 2025

#453

Interesting that the proofs seem to use a limited vocabulary: https://github.com/aw31/openai-imo-2025-proofs/blob/main/pro... Why waste time say lot word when few word do trick :) Also worth pointing out that Alex Wei is himself a gold medalist at IOI.

Interesting observation. One one hand, these resemble more the notes that an actual participant would write while solving the problem. Also, less words = less noise, more focus. But also, specifically for LLMs that output one token at a time and have a limited token context, I wonder if limiting itself to semantically meaningful tokens can be create longer stretches of semantically coherent thought?

Re: OpenAI claims gold-medal performance at IMO 2025

#454
I get the feeling that modern computer systems are so powerful that they can solve almost all well-explored closed problems with a properly tuned model. The problem lies in efficiency, reliability, and cost. Increasing efficiency and reliability would require an exponential increase in cost. QC might solve that cost part, and symbolic reasoning model will significantly boost both efficiency and reliability.

Re: OpenAI claims gold-medal performance at IMO 2025

#455
post #9

Some previous predictions: In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025. He thought there was an 8% chance of this happening. Eliezer Yudkowsky said "at least 16%". Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...

Off topic, but am I the only one getting triggered every time I see a rationalist quantify their prediction of the future with single digit accuracy? It's like their magic way of trying to get everyone to forget that they reached their conclusion in completely hand-wavy way, just like every other human being. But instead of saying "low confidence" or "high confidence" like the rest of us normies, they will tell you t…

If you take it with a grain of salt it's better than nothing. In life to express your opinion sometimes the best way is to quantify that based on intuition. To make decisions you could compile multiple experts intuitive quantities and use median or similar. There are some cases where it's more straight forward and rote, e.g. in military if you have to make distance based decisions, you might ask 8 of your soldiers to each name a number they think the distance is and take the median.

Re: OpenAI claims gold-medal performance at IMO 2025

#456
post #445

Earlier quoted context omitted.

I’ve been thinking a lot about what AI means about being human. Not about apocalypses or sentience or taking our jobs, but about “what is a human” and “what is the value of a human”. All my life I’ve taken for granted that your value is related to your positive impact, and that the unique value of humans is that we can create and express things like no other species we’ve encountered. Now, we have created this thing…

The fact that you're honestly grappling with this reality puts you far ahead of most people. It seems a common recent neurosis (albeit protective one) to proclaim a permanent human preeminence over the world of value, moral status and such for reasons extremely coupled with our intelligence, and then claim that certain kinds of intelligence have nothing to do with it when our primacy in those specific realms of intel…

That’s very insightful, thank you.

I agree that denial is not an approach that’s likely to be productive.

Re: OpenAI claims gold-medal performance at IMO 2025

#458

Earlier quoted context omitted.

Likely vs. unlikely is rounding to 50%. Single digit is rounding to 1%. I don't think the parent was suggesting the former is better than the latter. Even before I read your comment I thought that 5% precision is useful but 1% precision is a silly turn-off, unless that 1% is near the 0% or 100% boundary.

I thought single digit means single significant digit, aka rounding to 10%?

I did mean 1%, not sure if I used the right term though, english not being my first language.

Re: OpenAI claims gold-medal performance at IMO 2025

#459
post #252

Earlier quoted context omitted.

its honestly ruining this website, you cant even read the comments sections anymore

But in the case of OpenAI, this is fully justified. Isn't that so?

No. Even if it’s true, it adds nothing to the conversation and especially considering the article.
Post reply on HN