Noam Brown: > this isn’t an IMO-specific model. It’s a reasoning LLM that incorporates new experimental general-purpose techniques. > it’s also more efficient [than o1 or o3] with its thinking. And there’s a lot of room to push the test-time compute and efficiency further. > As fast as recent AI progress has been, I fully expect the trend to continue. Importantly, I think we’re close to AI substantially contributing…
> I think we’re close to AI substantially contributing to scientific discovery. The new "Full Self-Driving next year"?
OpenAI claims gold-medal performance at IMO 2025
471–480 of 737 posts
Re: OpenAI claims gold-medal performance at IMO 2025
#472I encourage anyone who thinks these are easy high-school problems to try to solve some. They're published (including this year's) at https://www.imo-official.org/problems.aspx . They make my head spin.
Re: OpenAI claims gold-medal performance at IMO 2025
#473Earlier quoted context omitted.
With regard to AI & LLMs Twitter/x is actually the only place with all of the industry people discussing. There are a bunch of great accounts to follow that are only really posting content to x. Karpathy, nearcyan, kalomaze, all of the OpenAI researchers including the link this discussion is on, many anthropic researchers. It's such a meme that you see people discuss reading Twitter thread + paper because the thread…
I'm unconvinced Twitter is a very good medium for serious technical discussion. I imagine a lot of this happens on the sidelines at conferences, on mailing threads and actually in organisations doing work on AI (e.g. Universities, Anthropic). The people who are doing the work are also often not the people who have time to Twitter.
Re: OpenAI claims gold-medal performance at IMO 2025
#474I encourage anyone who thinks these are easy high-school problems to try to solve some. They're published (including this year's) at https://www.imo-official.org/problems.aspx . They make my head spin.
- A 3Blue1Brown video on a particularly nice and unexpectedly difficult IMO problem (2011 IMO, Q2): https://www.youtube.com/watch?v=M64HUIJFTZM
-- And another similar one (though technically Putnam, not IMO): https://www.youtube.com/watch?v=OkmNXy7er84
- Timothy Gowers (Fields Medalist and IMO perfect scorer) solving this year’s IMO problems in “real time”:
Re: OpenAI claims gold-medal performance at IMO 2025
#475Earlier quoted context omitted.
> dismissing any claim from employees of large tech companies Me: I have a way to turn lead into gold. You: Show me!!! Me: NO (and then spends the rest of my life in poverty). Cold Fusion (physics not the programing language) is the best example of why you "Show your work". This is the Valley we're talking about. It's the thudnderdome of technology and companies. If you have a meaningful breakthrough you don't talk a…
I don't think this is a reasonable take. Some people/organizations send signals about things that we're not ready to fully drop it on the world. Others consider those signals in context (reputation of sender, prior probability of being true, reasons for sender to be honest vs. deceptive, etc). When my wife tells me there's a pie in the oven and it's smelling particularly good, I don't demand evidence or disbelieve th…
This is called marketing.
> When my wife tells me there's a pie in the oven and it's smelling particularly good, I don't demand evidence
Because you have evidence, it smells.
And if later your ask your wife "where is the pie" and she says "I sprayed pie scent in the air, I was just singling" how are you going to feel?
Open AI spent its "fool us once" card already. Doing things this way does not earn back trust, failure to deliver (and they have done that more than once) ... See staff non disparagement, see the math fiasco, see open weights.
Re: OpenAI claims gold-medal performance at IMO 2025
#476Earlier quoted context omitted.
I think the main hesitancy is due to rampant anthropomorphism. These models cannot reason, they pattern match language tokens and generate emergent behaviour as a result. Certainly the emergent behaviour is exciting but we tend to jump to conclusions as to what it implies. This means we are far more trusting with software that lacks formal guarantees than we should be. We are used to software being sound by default b…
> These models cannot reason Not trying to be a smarty pants here, but what do we mean by "reason"? Just to make the point, I'm using Claude to help me code right now. In between prompts, I read HN. It does things for me such as coding up new features, looking at the compile and runtime responses, and then correcting the code. All while I sit here and write with you on HN. It gives me feedback like "lock free message…
Re: OpenAI claims gold-medal performance at IMO 2025
#477Earlier quoted context omitted.
Yes, and there's also languages of ex-USSR countries, whose competitors presumably all understand Russian, and so on. The real reason might be that there's an enormous class of self-loathing elites in India who actively despise the possibility of any Indian language being represented in higher education. This obviously stunts the possibility of them being used in international competitions.
Discussions about Indian politics or the Indian psyche—especially when laced with Indic supremacist undertones—are off-topic and an annoyance here. Please consider sharing these views in a forum focused on Indian affairs, where they’re more likely to find the traction they deserve.
Discussions online have a tendency to go off into tangents like this. It's regrettable that this is such a contentious topic.
Re: OpenAI claims gold-medal performance at IMO 2025
#478Earlier quoted context omitted.
There is no honor in hiding behind euphemisms. Rationalists say ‘low confidence’ and ‘high confidence’ all the time, just not when they're making an actual bet and need to directly compare credences. And the 16.27% mockery is completely dishonest. They used less than a single significant figure.
> just not when they're making an actual bet That is not my experience talking with rationalists irl at all. And that is precisely my issue, it is pervasive in every day discussion about any topic, at least with the subset of rationalists I happen to cross paths with. If it was just for comparing ability to forecast or for bets, then sure it would make total sense. Just the other day I had a conversation with someone…
Re: OpenAI claims gold-medal performance at IMO 2025
#479Earlier quoted context omitted.
Off topic, but am I the only one getting triggered every time I see a rationalist quantify their prediction of the future with single digit accuracy? It's like their magic way of trying to get everyone to forget that they reached their conclusion in completely hand-wavy way, just like every other human being. But instead of saying "low confidence" or "high confidence" like the rest of us normies, they will tell you t…
Interestingly, this is actually a question that's been looked at empirically! Take a look at this paper: https://scholar.harvard.edu/files/rzeckhauser/files/value_of... They took high-precision forecasts from a forecasting tournament and rounded them to coarser buckets (nearest 5%, nearest 10%, nearest 33%), to see if the precision was actually conveying any real information. What they found is that if you rounded th…
Re: OpenAI claims gold-medal performance at IMO 2025
#480This is such an interesting time because the percentage of people who are making predictions about AGI happening on the future are going to drop off and the number of people completely ignoring the term AGI will increase.