>AI model performs astounding feat everyone claimed was impossible or won’t be achieved for a while >Commenters on HN claim it must not be that hard, or OpenAI is lying, or cheated. Anything but admit that it is impressive Every time on this site lol. A lot of people here have an emotional aversion to accepting AI progress. They’re deep in the bargaining/anger/denial phase.
I guess my major question would be: does the training data include anything from 2025 which may have included information about the IMO 2025? Given that AI companies are constantly trying to slurp up any and all data online, if the model was derived from existing work, it's maybe less impressive than at first glance. If present-day model does well at IMO 2026, that would be nice.
OpenAI claims gold-medal performance at IMO 2025
451–460 of 737 posts
Re: OpenAI claims gold-medal performance at IMO 2025
#452I tried P1 on chatgpt-o4-high, it tells me the solution is k=0 or 1. It doesn’t even know that k=3 is a solution for n=3. Such a solution would get 0/7 in the actual IMO.
Edit due to rate-limiting:
o3-pro returned an answer after 24 minutes: https://chatgpt.com/share/687bf8bf-c1b0-800b-b316-ca7dd9b009... Whether the CoT amounts to valid mathematical reasoning, I couldn't say, especially because OpenAI models tend to be very cagey with their CoT.
Gemini 2.5 Pro seems to have used more sophisticated reasoning ( https://g.co/gemini/share/c325915b5583 ) but it got a slightly different answer. Its chain of thought was unimpressive to say the least, so I'm not sure how it got its act together for the final explanation.
Claude Opus 4 appears to have missed the main solution the others found: https://claude.ai/share/3ba55811-8347-4637-a5f0-fd8790aa820b
Be interesting if someone could try Grok 4.
Re: OpenAI claims gold-medal performance at IMO 2025
#453Interesting that the proofs seem to use a limited vocabulary: https://github.com/aw31/openai-imo-2025-proofs/blob/main/pro... Why waste time say lot word when few word do trick :) Also worth pointing out that Alex Wei is himself a gold medalist at IOI.
Re: OpenAI claims gold-medal performance at IMO 2025
#454Re: OpenAI claims gold-medal performance at IMO 2025
#455Some previous predictions: In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025. He thought there was an 8% chance of this happening. Eliezer Yudkowsky said "at least 16%". Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...
Off topic, but am I the only one getting triggered every time I see a rationalist quantify their prediction of the future with single digit accuracy? It's like their magic way of trying to get everyone to forget that they reached their conclusion in completely hand-wavy way, just like every other human being. But instead of saying "low confidence" or "high confidence" like the rest of us normies, they will tell you t…
Re: OpenAI claims gold-medal performance at IMO 2025
#456Earlier quoted context omitted.
I’ve been thinking a lot about what AI means about being human. Not about apocalypses or sentience or taking our jobs, but about “what is a human” and “what is the value of a human”. All my life I’ve taken for granted that your value is related to your positive impact, and that the unique value of humans is that we can create and express things like no other species we’ve encountered. Now, we have created this thing…
The fact that you're honestly grappling with this reality puts you far ahead of most people. It seems a common recent neurosis (albeit protective one) to proclaim a permanent human preeminence over the world of value, moral status and such for reasons extremely coupled with our intelligence, and then claim that certain kinds of intelligence have nothing to do with it when our primacy in those specific realms of intel…
I agree that denial is not an approach that’s likely to be productive.
Re: OpenAI claims gold-medal performance at IMO 2025
#457I tried P1 on chatgpt-o4-high, it tells me the solution is k=0 or 1. It doesn’t even know that k=3 is a solution for n=3. Such a solution would get 0/7 in the actual IMO.
Re: OpenAI claims gold-medal performance at IMO 2025
#458Earlier quoted context omitted.
Likely vs. unlikely is rounding to 50%. Single digit is rounding to 1%. I don't think the parent was suggesting the former is better than the latter. Even before I read your comment I thought that 5% precision is useful but 1% precision is a silly turn-off, unless that 1% is near the 0% or 100% boundary.
I thought single digit means single significant digit, aka rounding to 10%?
Re: OpenAI claims gold-medal performance at IMO 2025
#459Earlier quoted context omitted.
its honestly ruining this website, you cant even read the comments sections anymore
But in the case of OpenAI, this is fully justified. Isn't that so?