Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

621–630 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#621

Earlier quoted context omitted.

because competitive coding is narrow well described domain(limited number of concepts: lists, trees, etc) with high volume of data available for training, and easy way to setup RL feeback loop, so models can improve well in this domain, which is not true about typical enterprise overbloated software.

All you said is true. Keep in mind this is the "Heuristics" competition instead of the "Algorithms" one. Instead of the more traditional Leetcode-like problems, it's things like optimizing scheduling/clustering according to some loss function. Think simulated annealing or pruned searches.

Dude thank you for stating this.

OpenAI's o3 model can solve very standard even up to 2700 rated codeforces problems it's been trained on, but is unable to think from first principles to solve problems I've set that are ~1600 rated. Those 2700 algorithms problems are obscure pages on the competitive programming wiki, so it's able to solve it with knowledge alone.

I am still not very impressed with its ability to reason both in codeforces and in software engineering. It's a very good database of information and a great searcher, but not a truly good first-principles reasoner.

I also wish o3 was a bit nicer - it's "reasoning" seems to have made it more arrogant at times too even when it's wildly off ,and it kind of annoys me.

Ironically, this workflow has really separated for me what is the core logic I should care about and what I should google, which is always a skill to learn when traversing new territory.

Re: OpenAI claims gold-medal performance at IMO 2025

#622

If someone told me this say, 10 or 20 years ago, I would have assumed this was worthy of a Nobel/Turing prize ...

Early machine learning researchers literally got Nobel Prize last year. Clearly not every incremental step of progress merits a Nobel.

Yes! But there is also a delay as 10 or 20 years ago we already had neural nets and I'm curious if people back then thought the concept was Nobel worthy.

Re: OpenAI claims gold-medal performance at IMO 2025

#623

Earlier quoted context omitted.

Meh. Some over hype, some under hype. People like you whine and then don't want to listen to any technical concerns. Some of us are implementing things in relation to AI so we know it's not about "increasing performance of models" but actual about the right solution for the right problem. If you think Twitter has "educated takes" then maybe go there and stop being pretentious schmuck over here. Talent drain, lol. I'd…

Both sides are not equally wrong, clearly. Until yesterday prediction markets were saying the probability of an AI getting a gold medal in IMO in 2025 was <20%. So clearly we should be more hyped, not less.

Prediction markets are companies and people trying to make money on volatility. Who cares? Why do people treat them as some prescient being?

Re: OpenAI claims gold-medal performance at IMO 2025

#624

Earlier quoted context omitted.

Thank you for this pop psychology evaluation. It could have been written by an an "AI".

Simply stating what is going on in the comments, as happens any time AI hits a previously thought to be impossible or far-off milestone.

I'm very much not in denial about my open source being being stolen without attribution or the layoffs that are rationalized by fake "AI" productivity.

You do have a point though that we should be writing to Sen. Marsha Blackburn instead of complaining here.

Re: OpenAI claims gold-medal performance at IMO 2025

#625
post #210

Earlier quoted context omitted.

> I've been reading this website for probably 15 years, its never been this bad. People here were pretty skeptical about AlexNet, when it won the ImageNet challenge 13 years ago. https://news.ycombinator.com/item?id=4611830

That thread was skeptical, but it's still far more substantive than what you find here today.

I think that's because the announcement there actually told you something technically interesting. This just presents a result (which is cool), but the actual method is what is really cool!

Re: OpenAI claims gold-medal performance at IMO 2025

#626
post #555

Earlier quoted context omitted.

Wasn't 16% the example they were talking about? Isn't that two significant digits? And 16% very much feels ridiculous to a reader when they could've just said 15%.

In context, the "at least 16%" is responding to someone who said 8%, and 16 just happens to be exactly twice 8. I suspect (though I don't know) that Yudkowsky would not have claimed to have a robust way to pick whether 16% or 17% was the better figure. For what it's worth, I don't think there's anything even slightly wrong with using whatever estimate feels good to you, even if it happens not to fit someone else's cr…

> In context, the "at least 16%" is responding to someone who said 8%, and 16 just happens to be exactly twice 8. I suspect (though I don't know) that Yudkowsky would not have claimed to have a robust way to pick whether 16% or 17% was the better figure.

If this was just a way to say "at least double that", that's... fair enough, I guess.

Regarding your other point:

> For what it's worth, I don't think there's anything even slightly wrong with using whatever estimate feels good to you, even if it happens not to fit someone else's criterion for being a nice round number

This is completely missing the point. There absolutely is something wrong with doing this (barring cases like the above where it was just a confusing phrasing of something with less precision like "double that"). The issue has nothing to do with being "nice", it has to do with the significant figures and the error bars.

If you say 20% then it is understood that your error margin is 5%. Even those that don't understand sigfigs still understand that your error margin is If you say 19% then suddenly the understanding becomes that your error margin Nobody is going to see that and assume your error bars on it are 5% -- nobody. Which is what makes it a ridiculous estimate. This has nothing to do with being "nice and round" and everything with conveying appropriate confidence.

Re: OpenAI claims gold-medal performance at IMO 2025

#627
post #9

Some previous predictions: In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025. He thought there was an 8% chance of this happening. Eliezer Yudkowsky said "at least 16%". Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...

One of the most worrying trends in AI has been how wrong the experts have been with overestimating timelines.

On the other hand, I think human hubris naturally makes us dramatically overestimate how special brains are.

Re: OpenAI claims gold-medal performance at IMO 2025

#628
post #85

I encourage anyone who thinks these are easy high-school problems to try to solve some. They're published (including this year's) at https://www.imo-official.org/problems.aspx . They make my head spin.

I like watching youtube videos solving these problems. They're deceptively simple. I remember reading one:

x+y=1

xy=1

The incredible thing is the explanation uses almost all reasoning steps that I am familiar with from basic algebra, like factoring, quadratic formula, etc. But it just comes together so beautifully. It gives you the impression that if you thought about it long enough, surely you would have come up with the answer, which is obviously wrong, at least in my case.

https://www.youtube.com/watch?v=csS4BjQuhCc

Re: OpenAI claims gold-medal performance at IMO 2025

#629
post #481

Earlier quoted context omitted.

Likely vs. unlikely is rounding to 50%. Single digit is rounding to 1%. I don't think the parent was suggesting the former is better than the latter. Even before I read your comment I thought that 5% precision is useful but 1% precision is a silly turn-off, unless that 1% is near the 0% or 100% boundary.

The book Superforecasting documented that for their best forecasters, rounding off that last percent would reliably reduce Brier scores. Whether rationalists who are publicly commenting actually achieve that level of reliability is an open question. But that humans can be reliable enough in the real world that the last percentage matters, has been demonstrated.

Your comment is incredibly confusing (possibly misleading) because of the key details you've omitted.

> The book Superforecasting documented that for their best forecasters, rounding off that last percent would reliably reduce Brier scores.

Rounding off that last percent... to what, exactly? Are you excluding the exceptions I mentioned (i.e. when you're already close to 0% or 100%?)

Nobody is arguing that 3% -> 4% is insignificant. The argument is over whether 16% -> 15% is significant.

Re: OpenAI claims gold-medal performance at IMO 2025

#630

Earlier quoted context omitted.

> We can only go off their word We’re talking about Sam Altman’s company here. The same company that started out as a non profit claiming they wanted to better the world. Suggesting they should be given the benefit of the doubt is dishonest at this point.

“they must be lying because I personally dislike them” This is why HN threads about AI have become exhausting to read

No, they are likely lying, because they have huge incentives to lie
Post reply on HN