Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

131–140 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#131
post #96

OpenAI simply can’t be trusted on any benchmarks: https://news.ycombinator.com/item?id=42761648

Remember that they've fired all whistleblowers that would admit to breaking the verbal agreement that they wouldn't train on the test data.

Could not find it on the open web. Do you have clues to search for?

Re: OpenAI claims gold-medal performance at IMO 2025

#132

I think equally impressive is the performance of the OpenAI team at the "AtCoder World Tour Finals 2025" a couple of days ago. There were 12 human participants and only one did better than OpenAI. Not sure there is a good writeup about it yet but here is the livestream: https://www.youtube.com/live/TG3ChQH61vE .

And yet when working on production code current LLMs are about as good as a poor intern. Not sure why the disconnect.

Re: OpenAI claims gold-medal performance at IMO 2025

#133
The cynicism/denial on HN about AI is exhausting. Half the comments are some weird form of explaining away the ever increasing performance of these models

I've been reading this website for probably 15 years, its never been this bad. many threads are completely unreadable, all the actual educated takes are on X, its almost like there was a talent drain

Re: OpenAI claims gold-medal performance at IMO 2025

#134

My issue with all these citations is that it’s all OpenAI employees that make these claims. I’ll wait to see third party verification and/or use it myself before judging. There’s a lot of incentives right now to hype things up for OpenAI.

A third party tried this experiment with publicly available models. OpenAI did half as well as Gemini, and none of the models even got bronze.

https://matharena.ai/imo/

Re: OpenAI claims gold-medal performance at IMO 2025

#135
post #86

Earlier quoted context omitted.

The key bit here is whether the LLM doing the cherry picking had knowledge of the solution. If it didn't, this is a meaningful result. That's why I'd like more info, but I fear OpenAI is going to try to keep things under wraps.

> If it didn't We kind of have to assume it didn't right? Otherwise bragging about the results makes zero sense and would be outright misleading.

openai have been caught doing exactly this before

Re: OpenAI claims gold-medal performance at IMO 2025

#136

The cynicism/denial on HN about AI is exhausting. Half the comments are some weird form of explaining away the ever increasing performance of these models I've been reading this website for probably 15 years, its never been this bad. many threads are completely unreadable, all the actual educated takes are on X, its almost like there was a talent drain

Probably because both sides have strong vested interests and it’s next to impossible to find a dispassionate point of view.

The Pro AI crowd, VC, tech CEOs etc have strong incentive to claim humans are obsolete. Many tech employees see threats to their jobs and want to poopoo any way AI could be useful or competitive.

Re: OpenAI claims gold-medal performance at IMO 2025

#137
The issue is that trust is very hard to build and very easy to lose. Even in today's age where regular humans have a memory span shorter than that of an LLM, OpenAI keeps abusing the public's trust. As a result, I take their word on AI/LLMs about as seriously as I'd take my grocery store clerk's opinion on quantum physics.

Re: OpenAI claims gold-medal performance at IMO 2025

#139
post #82
post #63

Earlier quoted context omitted.

I feel like I've noticed you you making the same comment 12 places in this thread -- incorrectly misrepresenting the difficulty of this tournament and ultimately it comes across as a bitter ex. Here's an example problem 5: Let a1,a2,…,an be distinct positive integers and let M=max⁡1≤i Find the maximum number of pairs (i,j) with 1≤i<j≤n for which (ai +aj )(aj −ai )=M.

What does max⁡1≤i<j≤n mean? Wouldn't M always be j?

[deleted]

Re: OpenAI claims gold-medal performance at IMO 2025

#140

The cynicism/denial on HN about AI is exhausting. Half the comments are some weird form of explaining away the ever increasing performance of these models I've been reading this website for probably 15 years, its never been this bad. many threads are completely unreadable, all the actual educated takes are on X, its almost like there was a talent drain

The overconfidence/short sightedness on HN about AI is exhausting. Half the comments are some weird form of explaining how developers will be obsolete in five years and how close we are to AGI.
Post reply on HN