> level performance on the world’s most prestigious math competition I don't know which one i would consider the most prestigious math competition but it wouldn't be The IMO. The Putnam ranks higher to me and I'm not even an American. But I've come to realise one thing and that is that high-school is very important to Americans...
OpenAI claims gold-medal performance at IMO 2025
461–470 of 737 posts
Re: OpenAI claims gold-medal performance at IMO 2025
#462Earlier quoted context omitted.
Off topic, but am I the only one getting triggered every time I see a rationalist quantify their prediction of the future with single digit accuracy? It's like their magic way of trying to get everyone to forget that they reached their conclusion in completely hand-wavy way, just like every other human being. But instead of saying "low confidence" or "high confidence" like the rest of us normies, they will tell you t…
No, you are right, this hyper-numericalism is just astrology for nerds.
Re: OpenAI claims gold-medal performance at IMO 2025
#463Earlier quoted context omitted.
> I've been reading this website for probably 15 years, its never been this bad... all the actual educated takes are on X Almost every technical comment on HN is wrong (see for example essentially all the discussion of Rust async, in which people keep making up silly claims that Rust maintainers then attempt to patiently explain are wrong). The idea that the "educated" takes are on X though... that's crazy talk.
With regard to AI & LLMs Twitter/x is actually the only place with all of the industry people discussing. There are a bunch of great accounts to follow that are only really posting content to x. Karpathy, nearcyan, kalomaze, all of the OpenAI researchers including the link this discussion is on, many anthropic researchers. It's such a meme that you see people discuss reading Twitter thread + paper because the thread…
Re: OpenAI claims gold-medal performance at IMO 2025
#464Some previous predictions: In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025. He thought there was an 8% chance of this happening. Eliezer Yudkowsky said "at least 16%". Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...
Off topic, but am I the only one getting triggered every time I see a rationalist quantify their prediction of the future with single digit accuracy? It's like their magic way of trying to get everyone to forget that they reached their conclusion in completely hand-wavy way, just like every other human being. But instead of saying "low confidence" or "high confidence" like the rest of us normies, they will tell you t…
And since we’re at it: why not give confidence intervals too?
Re: OpenAI claims gold-medal performance at IMO 2025
#465Earlier quoted context omitted.
I think the main hesitancy is due to rampant anthropomorphism. These models cannot reason, they pattern match language tokens and generate emergent behaviour as a result. Certainly the emergent behaviour is exciting but we tend to jump to conclusions as to what it implies. This means we are far more trusting with software that lacks formal guarantees than we should be. We are used to software being sound by default b…
> These models cannot reason Not trying to be a smarty pants here, but what do we mean by "reason"? Just to make the point, I'm using Claude to help me code right now. In between prompts, I read HN. It does things for me such as coding up new features, looking at the compile and runtime responses, and then correcting the code. All while I sit here and write with you on HN. It gives me feedback like "lock free message…
Re: OpenAI claims gold-medal performance at IMO 2025
#466I think equally impressive is the performance of the OpenAI team at the "AtCoder World Tour Finals 2025" a couple of days ago. There were 12 human participants and only one did better than OpenAI. Not sure there is a good writeup about it yet but here is the livestream: https://www.youtube.com/live/TG3ChQH61vE .
And yet when working on production code current LLMs are about as good as a poor intern. Not sure why the disconnect.
Re: OpenAI claims gold-medal performance at IMO 2025
#467Earlier quoted context omitted.
While I usually enjoy seeing these discussions, I think they are really pushing the usefulness of bayesian statistics. If one dude says the chance for an outcome is 8% and another says it's 16% and the outcome does occur, they were both pretty wrong, even though it might seem like the one who guessed a few % higher might have had a better belief system. Now if one of them had said 90% while the other said 8% or 16%,…
The correctness of 8%, 16%, and 90% are all equally unknown since we only have one timeline, no?
Re: OpenAI claims gold-medal performance at IMO 2025
#468Some previous predictions: In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025. He thought there was an 8% chance of this happening. Eliezer Yudkowsky said "at least 16%". Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...
Context? Who are these people and what are these numbers and why shouldn't I assume they're pulled from thin air?
Clowns, mostly. Yudkowski in particular, whose only job today seems to be making awful predictions and letting lesswrong eat it up when one out of a hundred ends up coming true, solidifying his position as AI-will-destroy-the-world messiah. They make money from these outlandish takes, and more money when you keep talking about them.
It's kind of like listening to the local drunkard at the bar that once in a while ends up predicting which team is going to win in football inbetween drunken and nonsensical rants, except that for some reason posting the predictions on the internet makes him a celebrity, instead of just a drunk curiosity.
Re: OpenAI claims gold-medal performance at IMO 2025
#469Earlier quoted context omitted.
Likely vs. unlikely is rounding to 50%. Single digit is rounding to 1%. I don't think the parent was suggesting the former is better than the latter. Even before I read your comment I thought that 5% precision is useful but 1% precision is a silly turn-off, unless that 1% is near the 0% or 100% boundary.
I thought single digit means single significant digit, aka rounding to 10%?
And 16% very much feels ridiculous to a reader when they could've just said 15%.
Re: OpenAI claims gold-medal performance at IMO 2025
#470Earlier quoted context omitted.
Making an account just to point out how these comments are far more exhausting, because they don't engage with the subject matter. They are just agreeing with a headline and saying, "See?" You say, "explaining away the increasing performance" as though that was a good faith representation of arguments made against LLMs, or even this specific article. Questionong the self-congragulatory nature of these businesses is p…
But don't you think this might be a case where there is both self-congragulation and actual progress?