OpenAI simply can’t be trusted on any benchmarks: https://news.ycombinator.com/item?id=42761648
Remember that they've fired all whistleblowers that would admit to breaking the verbal agreement that they wouldn't train on the test data.
OpenAI claims gold-medal performance at IMO 2025
131–140 of 737 posts
Re: OpenAI claims gold-medal performance at IMO 2025
#132I think equally impressive is the performance of the OpenAI team at the "AtCoder World Tour Finals 2025" a couple of days ago. There were 12 human participants and only one did better than OpenAI. Not sure there is a good writeup about it yet but here is the livestream: https://www.youtube.com/live/TG3ChQH61vE .
Re: OpenAI claims gold-medal performance at IMO 2025
#133I've been reading this website for probably 15 years, its never been this bad. many threads are completely unreadable, all the actual educated takes are on X, its almost like there was a talent drain
Re: OpenAI claims gold-medal performance at IMO 2025
#134My issue with all these citations is that it’s all OpenAI employees that make these claims. I’ll wait to see third party verification and/or use it myself before judging. There’s a lot of incentives right now to hype things up for OpenAI.
Re: OpenAI claims gold-medal performance at IMO 2025
#135Earlier quoted context omitted.
The key bit here is whether the LLM doing the cherry picking had knowledge of the solution. If it didn't, this is a meaningful result. That's why I'd like more info, but I fear OpenAI is going to try to keep things under wraps.
> If it didn't We kind of have to assume it didn't right? Otherwise bragging about the results makes zero sense and would be outright misleading.
Re: OpenAI claims gold-medal performance at IMO 2025
#136The cynicism/denial on HN about AI is exhausting. Half the comments are some weird form of explaining away the ever increasing performance of these models I've been reading this website for probably 15 years, its never been this bad. many threads are completely unreadable, all the actual educated takes are on X, its almost like there was a talent drain
The Pro AI crowd, VC, tech CEOs etc have strong incentive to claim humans are obsolete. Many tech employees see threats to their jobs and want to poopoo any way AI could be useful or competitive.
Re: OpenAI claims gold-medal performance at IMO 2025
#137Re: OpenAI claims gold-medal performance at IMO 2025
#138I believe this company used to present its results and approach in academic papers with enough details so that it could be reproduced by third parties. Now it is just doing a bunch of tweets?
Re: OpenAI claims gold-medal performance at IMO 2025
#139Earlier quoted context omitted.
I feel like I've noticed you you making the same comment 12 places in this thread -- incorrectly misrepresenting the difficulty of this tournament and ultimately it comes across as a bitter ex. Here's an example problem 5: Let a1,a2,…,an be distinct positive integers and let M=max1≤i Find the maximum number of pairs (i,j) with 1≤i<j≤n for which (ai +aj )(aj −ai )=M.
What does max1≤i<j≤n mean? Wouldn't M always be j?
Re: OpenAI claims gold-medal performance at IMO 2025
#140The cynicism/denial on HN about AI is exhausting. Half the comments are some weird form of explaining away the ever increasing performance of these models I've been reading this website for probably 15 years, its never been this bad. many threads are completely unreadable, all the actual educated takes are on X, its almost like there was a talent drain