Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

251–260 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#251

The cynicism/denial on HN about AI is exhausting. Half the comments are some weird form of explaining away the ever increasing performance of these models I've been reading this website for probably 15 years, its never been this bad. many threads are completely unreadable, all the actual educated takes are on X, its almost like there was a talent drain

Software that mangles data on the regular should be thrown away.

How is it rational to 10x the budget over and over again when it mangles data every time?

The mind blowing thing is not being skeptical of that approach, it's defending it. It has become an article of faith.

It would be great to have AI chatbots. But chatbots that mangle data getting their budgets increased by orders of magnitude over and over again is just doubling down on the same mistake over and over again.

Re: OpenAI claims gold-medal performance at IMO 2025

#252

Earlier quoted context omitted.

I don’t typically find this to be true. There is a definite cynicism on HN especially when it comes to OpenAI. You already know what you will see. Low quality garbage of “I remember when OpenAI was open”, “remember when they used to publish research”, “sama cannot be trusted”, it’s an endless barrage of garbage.

its honestly ruining this website, you cant even read the comments sections anymore

But in the case of OpenAI, this is fully justified. Isn't that so?

Re: OpenAI claims gold-medal performance at IMO 2025

#253
post #92

In the RLHF sphere you could tell some AI company/companies were targeting this because of how many IMO RLHF’ers they were hiring specifically. I don’t think it’s really easy to say how much “progress” this is given that.

I doubt this is coming from RLHF - tweets from the lead researcher state that this result flows from a research breakthrough which enables RLVR on less verifiable domains.

Re: OpenAI claims gold-medal performance at IMO 2025

#254

Its a level playing field IMO. But theres another thread which claims not even bronze and I really don't want to go to X for anything.

I can save you the click. Public models (gemini/o3) are less than bronze. this is a specially trained model which is not publicly available.

Is the conclusion that elite models are being withheld from the public, or that general models are not that general?

Re: OpenAI claims gold-medal performance at IMO 2025

#255
post #89

Am I missing something or is this completely meaningless? It's 100% opaque, no details whatsoever and no transparency or reproducibility. I wouldn't trust these results as it is. Considering that there are trillions of dollars on the line as a reward for hyping up LLMs, I trust it even less.

Yes you are missing the entire boat

The entire boat is hidden, we see nothing but a projected shadow, but we are to be blamed for missing it?

Re: OpenAI claims gold-medal performance at IMO 2025

#256
post #9

Some previous predictions: In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025. He thought there was an 8% chance of this happening. Eliezer Yudkowsky said "at least 16%". Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...

Off topic, but am I the only one getting triggered every time I see a rationalist quantify their prediction of the future with single digit accuracy? It's like their magic way of trying to get everyone to forget that they reached their conclusion in completely hand-wavy way, just like every other human being. But instead of saying "low confidence" or "high confidence" like the rest of us normies, they will tell you they think there is 16.27% chance because they really really want you to be aware that they know bayes theorem.

Re: OpenAI claims gold-medal performance at IMO 2025

#257

The cynicism/denial on HN about AI is exhausting. Half the comments are some weird form of explaining away the ever increasing performance of these models I've been reading this website for probably 15 years, its never been this bad. many threads are completely unreadable, all the actual educated takes are on X, its almost like there was a talent drain

Enthusiastically denouncing or promoting something is much, much easier and more rewarding in the short term for people who want to appear hip to their chosen in-group - or profit center.

And then, it's likewise easy to be a reactionary to the extremes of the other side.

The middle is a harder, more interesting place to be, and people who end up there aren't usually chasing money or power, but some approximation of the truth.

Re: OpenAI claims gold-medal performance at IMO 2025

#258
post #91
post #9

Some previous predictions: In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025. He thought there was an 8% chance of this happening. Eliezer Yudkowsky said "at least 16%". Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...

Context? Who are these people and what are these numbers and why shouldn't I assume they're pulled from thin air?

> why shouldn't I assume they're pulled from thin air?

You definitely should assume they are. They are rationalists, the modus operandi is to pull stuff out of thin air and slap a single digit precision percentage prediction in front to make it seems grounded in science and well thought out.

Re: OpenAI claims gold-medal performance at IMO 2025

#260
post #57

Earlier quoted context omitted.

It's like saying getting a gold medal in boxing is not hard, because it doesn't involve any firearms

More fair comparison: Military grade killbot enters ring with boxer and proceeds to fire pneumatic hammer at boxer until KO?

So far reasoning models haven't been strong at, well, reasoning. If they are really like a killbot against a boxer now, that's pretty newsworthy
Post reply on HN