Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

461–470 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#461

> level performance on the world’s most prestigious math competition I don't know which one i would consider the most prestigious math competition but it wouldn't be The IMO. The Putnam ranks higher to me and I'm not even an American. But I've come to realise one thing and that is that high-school is very important to Americans...

The Putnam and IMO are quite different. I would suggest the IMO is probably harder...

Re: OpenAI claims gold-medal performance at IMO 2025

#462

Earlier quoted context omitted.

Off topic, but am I the only one getting triggered every time I see a rationalist quantify their prediction of the future with single digit accuracy? It's like their magic way of trying to get everyone to forget that they reached their conclusion in completely hand-wavy way, just like every other human being. But instead of saying "low confidence" or "high confidence" like the rest of us normies, they will tell you t…

No, you are right, this hyper-numericalism is just astrology for nerds.

In military they estimate distances this way if they don't have proper tools. Each says a min max range and then where there's most overlap, that will be taken. It's a reasonable way to make quick intuition based decisions when no other way is available.

Re: OpenAI claims gold-medal performance at IMO 2025

#463
post #215

Earlier quoted context omitted.

> I've been reading this website for probably 15 years, its never been this bad... all the actual educated takes are on X Almost every technical comment on HN is wrong (see for example essentially all the discussion of Rust async, in which people keep making up silly claims that Rust maintainers then attempt to patiently explain are wrong). The idea that the "educated" takes are on X though... that's crazy talk.

With regard to AI & LLMs Twitter/x is actually the only place with all of the industry people discussing. There are a bunch of great accounts to follow that are only really posting content to x. Karpathy, nearcyan, kalomaze, all of the OpenAI researchers including the link this discussion is on, many anthropic researchers. It's such a meme that you see people discuss reading Twitter thread + paper because the thread…

I'm unconvinced Twitter is a very good medium for serious technical discussion. I imagine a lot of this happens on the sidelines at conferences, on mailing threads and actually in organisations doing work on AI (e.g. Universities, Anthropic). The people who are doing the work are also often not the people who have time to Twitter.

Re: OpenAI claims gold-medal performance at IMO 2025

#464
post #9

Some previous predictions: In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025. He thought there was an 8% chance of this happening. Eliezer Yudkowsky said "at least 16%". Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...

Off topic, but am I the only one getting triggered every time I see a rationalist quantify their prediction of the future with single digit accuracy? It's like their magic way of trying to get everyone to forget that they reached their conclusion in completely hand-wavy way, just like every other human being. But instead of saying "low confidence" or "high confidence" like the rest of us normies, they will tell you t…

No you’re definitely not the only one… 10% is ok, 5% maybe, 1% is useless.

And since we’re at it: why not give confidence intervals too?

Re: OpenAI claims gold-medal performance at IMO 2025

#465

Earlier quoted context omitted.

I think the main hesitancy is due to rampant anthropomorphism. These models cannot reason, they pattern match language tokens and generate emergent behaviour as a result. Certainly the emergent behaviour is exciting but we tend to jump to conclusions as to what it implies. This means we are far more trusting with software that lacks formal guarantees than we should be. We are used to software being sound by default b…

> These models cannot reason Not trying to be a smarty pants here, but what do we mean by "reason"? Just to make the point, I'm using Claude to help me code right now. In between prompts, I read HN. It does things for me such as coding up new features, looking at the compile and runtime responses, and then correcting the code. All while I sit here and write with you on HN. It gives me feedback like "lock free message…

YOU are reasoning.

Re: OpenAI claims gold-medal performance at IMO 2025

#466

I think equally impressive is the performance of the OpenAI team at the "AtCoder World Tour Finals 2025" a couple of days ago. There were 12 human participants and only one did better than OpenAI. Not sure there is a good writeup about it yet but here is the livestream: https://www.youtube.com/live/TG3ChQH61vE .

And yet when working on production code current LLMs are about as good as a poor intern. Not sure why the disconnect.

It’s the same reason leet code is a bad interview question. Being good at these sorts of problems doesn’t translate directly to being good at writing production code.

Re: OpenAI claims gold-medal performance at IMO 2025

#467

Earlier quoted context omitted.

While I usually enjoy seeing these discussions, I think they are really pushing the usefulness of bayesian statistics. If one dude says the chance for an outcome is 8% and another says it's 16% and the outcome does occur, they were both pretty wrong, even though it might seem like the one who guessed a few % higher might have had a better belief system. Now if one of them had said 90% while the other said 8% or 16%,…

The correctness of 8%, 16%, and 90% are all equally unknown since we only have one timeline, no?

This is probably the best thing I’ve ever read about predictions of the future. If we could run 80 parallel universes then sure it would make sense. But we only have the one [1]. If you’re right and we get fast takeoff it won’t matter because we’re all dead. In any case the number is meaningless, there is only ONE future.

Re: OpenAI claims gold-medal performance at IMO 2025

#468
post #91
post #9

Some previous predictions: In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025. He thought there was an 8% chance of this happening. Eliezer Yudkowsky said "at least 16%". Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...

Context? Who are these people and what are these numbers and why shouldn't I assume they're pulled from thin air?

>Who are these people

Clowns, mostly. Yudkowski in particular, whose only job today seems to be making awful predictions and letting lesswrong eat it up when one out of a hundred ends up coming true, solidifying his position as AI-will-destroy-the-world messiah. They make money from these outlandish takes, and more money when you keep talking about them.

It's kind of like listening to the local drunkard at the bar that once in a while ends up predicting which team is going to win in football inbetween drunken and nonsensical rants, except that for some reason posting the predictions on the internet makes him a celebrity, instead of just a drunk curiosity.

Re: OpenAI claims gold-medal performance at IMO 2025

#469

Earlier quoted context omitted.

Likely vs. unlikely is rounding to 50%. Single digit is rounding to 1%. I don't think the parent was suggesting the former is better than the latter. Even before I read your comment I thought that 5% precision is useful but 1% precision is a silly turn-off, unless that 1% is near the 0% or 100% boundary.

I thought single digit means single significant digit, aka rounding to 10%?

Wasn't 16% the example they were talking about? Isn't that two significant digits?

And 16% very much feels ridiculous to a reader when they could've just said 15%.

Re: OpenAI claims gold-medal performance at IMO 2025

#470
post #267

Earlier quoted context omitted.

Making an account just to point out how these comments are far more exhausting, because they don't engage with the subject matter. They are just agreeing with a headline and saying, "See?" You say, "explaining away the increasing performance" as though that was a good faith representation of arguments made against LLMs, or even this specific article. Questionong the self-congragulatory nature of these businesses is p…

But don't you think this might be a case where there is both self-congragulation and actual progress?

That's a fair question, and I agree. I just find it odd how we shout across the aisle, whether in favor or against. It's a case of thinking the tech is neat, while cringing at all the money-people and their ideations.
Post reply on HN