Earlier quoted context omitted.
They also said this is not part of GPT-5, and “will be released later”. It’s very, very likely a model specifically fine-tuned for this benchmark, where afterwards they’ll evaluate what actual real-world problems it’s good at (eg like “use o4-mini-high for coding”).
Humans who excel at IMO questions are also "fine tuned" on them in the sense that they practice them for hundreds of hours
OpenAI claims gold-medal performance at IMO 2025
531–540 of 737 posts
Re: OpenAI claims gold-medal performance at IMO 2025
#532I am neither an optimist nor a pessimist for AI. I would likely be called both by the opposite parties. But the fact that AI / LLM is still rapidly improving is impressive in itself and worth celebrating for. Is it perfect, AGI, ASI? No. Is it useless? Absolutely not. I am just happy the prize is so big for AI that there are enough money involve to push for all the hardware advancement. Foundry, Packaging, Interconne…
The correlation between "companies make smarter AI" and "our lives get better" is still a rounding error.
Many people will say "don't worry, tech always makes our lives better eventually", they'll probably stop saying this once autonomous killer drone-swarms are a thing.
Re: OpenAI claims gold-medal performance at IMO 2025
#533Earlier quoted context omitted.
>I think the main hesitancy is due to rampant anthropomorphism. These models cannot reason, they pattern match language tokens and generate emergent behaviour as a result This is rampant human chauvinism. There's absolutely no empirical basis for the statement that these models "cannot reason", it's just pseudoscientific woo thrown around by people who want to feel that humans are somehow special. By pretty much ever…
I’ve used these AI tools for multiple hours a day for months. Not seeing the reasoning party honestly. I see the heuristics part.
Re: OpenAI claims gold-medal performance at IMO 2025
#534I am neither an optimist nor a pessimist for AI. I would likely be called both by the opposite parties. But the fact that AI / LLM is still rapidly improving is impressive in itself and worth celebrating for. Is it perfect, AGI, ASI? No. Is it useless? Absolutely not. I am just happy the prize is so big for AI that there are enough money involve to push for all the hardware advancement. Foundry, Packaging, Interconne…
But unlike the trillion dollars invested in the broadband internet build out between 1998 and 2008, when this 10 year trillion dollar bubble pops, we won't be left with an enduring and useful piece of infrastructure adding a trillion dollars to the global economy annually.
Re: OpenAI claims gold-medal performance at IMO 2025
#535Re: OpenAI claims gold-medal performance at IMO 2025
#536Earlier quoted context omitted.
Discussions about Indian politics or the Indian psyche—especially when laced with Indic supremacist undertones—are off-topic and an annoyance here. Please consider sharing these views in a forum focused on Indian affairs, where they’re more likely to find the traction they deserve.
It is not "supremacist" to believe that depriving hundreds of millions of people from higher education in their native language is deeply unjust. This reflection was prompted by a comment on why Indian languages are not represented in international competitions, which was prompted by a comment on the competition being available in many languages. Discussions online have a tendency to go off into tangents like this. I…
Re: OpenAI claims gold-medal performance at IMO 2025
#537Earlier quoted context omitted.
> such announcements should wait at least a week after the closing ceremony it would raise more concerns that corps leaked questions/answers to training data and finetuned specialized models during this time.
This is solvable? They could post a hash of the solutions publicly, and reveal the contents of the solution after one week.
Re: OpenAI claims gold-medal performance at IMO 2025
#538Earlier quoted context omitted.
>I think the main hesitancy is due to rampant anthropomorphism. These models cannot reason, they pattern match language tokens and generate emergent behaviour as a result This is rampant human chauvinism. There's absolutely no empirical basis for the statement that these models "cannot reason", it's just pseudoscientific woo thrown around by people who want to feel that humans are somehow special. By pretty much ever…
> This is rampant human chauvinism What in the accelerationist hell?
Re: OpenAI claims gold-medal performance at IMO 2025
#539Earlier quoted context omitted.
> and not so much that fact that they used precise numbers in the their internal monologue as opposed to verbal buckets like "pretty likely", "very unlikely" I am obviously only talking from my personal anecdotal experience, but having been on a bunch of coffee chat in the last few months with people in the AI safety field in SF, and a lot of them being Lesswrong-ers, I experienced a lot of those discussions with ran…
> It could be just a personal mental weakness with numbers with me that is not general, but looking at my interlocutors emotional reactions to their own numerical predictions I do feel quite strongly that this is a general human trait. Your feeling is correct; anchoring is a thing, and good LessWrongers (I hope to be in that category) know this and keep track of where their prior and not just posterior probabilities…
Re: OpenAI claims gold-medal performance at IMO 2025
#540Some previous predictions: In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025. He thought there was an 8% chance of this happening. Eliezer Yudkowsky said "at least 16%". Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...
Off topic, but am I the only one getting triggered every time I see a rationalist quantify their prediction of the future with single digit accuracy? It's like their magic way of trying to get everyone to forget that they reached their conclusion in completely hand-wavy way, just like every other human being. But instead of saying "low confidence" or "high confidence" like the rest of us normies, they will tell you t…
The rest of the sentence is not necessary. No, you're not the only one.