भारत दुर्दशा न देखी जाई...
OpenAI claims gold-medal performance at IMO 2025
381–390 of 737 posts
Re: OpenAI claims gold-medal performance at IMO 2025
#382Earlier quoted context omitted.
> We can only go off their word We’re talking about Sam Altman’s company here. The same company that started out as a non profit claiming they wanted to better the world. Suggesting they should be given the benefit of the doubt is dishonest at this point.
“they must be lying because I personally dislike them” This is why HN threads about AI have become exhausting to read
Those models seem to be special and not part of their normal product line, as is pointed out in the comments here. I would assume that in that case they indeed had the purpose of passing these tests in mind when creating them. Or was it created for something different, and completely by chance they discovered they could be used for the challenge, unintentionally?
Re: OpenAI claims gold-medal performance at IMO 2025
#383Noam Brown: > this isn’t an IMO-specific model. It’s a reasoning LLM that incorporates new experimental general-purpose techniques. > it’s also more efficient [than o1 or o3] with its thinking. And there’s a lot of room to push the test-time compute and efficiency further. > As fast as recent AI progress has been, I fully expect the trend to continue. Importantly, I think we’re close to AI substantially contributing…
That's a big leap from "answering test questions" to "contributing to scientific discovery".
Re: OpenAI claims gold-medal performance at IMO 2025
#384Noam Brown: > this isn’t an IMO-specific model. It’s a reasoning LLM that incorporates new experimental general-purpose techniques. > it’s also more efficient [than o1 or o3] with its thinking. And there’s a lot of room to push the test-time compute and efficiency further. > As fast as recent AI progress has been, I fully expect the trend to continue. Importantly, I think we’re close to AI substantially contributing…
> I think we’re close to AI substantially contributing to scientific discovery. The new "Full Self-Driving next year"?
Re: OpenAI claims gold-medal performance at IMO 2025
#385Earlier quoted context omitted.
Off topic, but am I the only one getting triggered every time I see a rationalist quantify their prediction of the future with single digit accuracy? It's like their magic way of trying to get everyone to forget that they reached their conclusion in completely hand-wavy way, just like every other human being. But instead of saying "low confidence" or "high confidence" like the rest of us normies, they will tell you t…
There is no honor in hiding behind euphemisms. Rationalists say ‘low confidence’ and ‘high confidence’ all the time, just not when they're making an actual bet and need to directly compare credences. And the 16.27% mockery is completely dishonest. They used less than a single significant figure.
That is not my experience talking with rationalists irl at all. And that is precisely my issue, it is pervasive in every day discussion about any topic, at least with the subset of rationalists I happen to cross paths with. If it was just for comparing ability to forecast or for bets, then sure it would make total sense.
Just the other day I had a conversation with someone about working in AI safety, it when something like "well I think there is 10 to 15% chance of AGI going wrong, and if I join I have maybe 1% chance of being able to make an impact and if.. and if... and if, so if we compare with what I'm missing by not going to instead I have 35% confidence it's the right decision"
What makes me uncomfortable with this, is that by using this kind of reasoning and coming out with a precise figure at the end, it cognitively bias you into being more confident in your reasoning than you should be. Because we are all used to treat numbers as the output of a deterministic, precise, scientific process.
There is no reason to say 10% or 15% and not 8% or 20% for rogue AGI, there is no reason to think one individual can change the direction by 1% and not by 0.3% or 3%, it's all just random numbers, and so when you multiply a gut feeling number by a gut feeling number 5 times in a row, you end up with something absolutely meaningless, where the margin of error is basically 100%.
But it somehow feels more scientific and reliable because it's a precise number, and I think this is dishonest and misleading both to the speaker themselves and to listeners. "Low confidence", or "im really not sure but I think..." have the merit of not hiding a gut feeling process behind a scientific veil.
To be clear I'm not saying you should never use numerics to try to quantify gut feeling, it's ok to say I think there is maybe 10% chance of rogue AGI and thus I want to do this or that. What I really don't like is the stacking of multiple random predictions and trying to reason about this in good faith.
> And the 16.27% mockery is completely dishonest.
Obviously satire
Re: OpenAI claims gold-medal performance at IMO 2025
#386Earlier quoted context omitted.
Implying results are fraudulent is completely fair when it is a fraud. The previous time they had claims about solving all of the math right there and right then, they were caught owning the company that makes that independent test , and could neither admit nor deny training on closed test set.
Just to quickly clarify: - OpenAI doesn't own Epoch AI (though they did commission Epoch to make the eval) - OpenAI denied training on the test set (and further denied training on FrontierMath-derived data, training on data targeting FrontierMath specifically, or using the eval to pick a model checkpoint; in fact, they only downloaded the FrontierMath data after their o3 training set was frozen and they didn't look a…
Everywhere I worked offered me a significant amount of money to sign a non-disparagement agreement after I left. I have never met someone who didn't willingly sign these agreements. The companies always make it clear if you refuse to sign they will give you a bad recommendation in the future.
Re: OpenAI claims gold-medal performance at IMO 2025
#387It’s interesting that this is a competition elite enough that several posters on a programming website don’t seem to understand what it is. My very rough napkin math suggests that against the US reference class, imo gold is literally a one in a million talent (very roughly 20 people who make camp could get gold out of very roughly twenty million relevant high schoolers).
I’m not trying to take away from the difficulty of the competition. But I went to a relatively well regarded high school and never even heard of IMO until I met competitors during undergrad. I think that the number of students who are even aware of the competition is way lower than the total number of students. I mean, I don’t think I’d have been a great competitor even if I tried. But I’m pretty sure there are a lot…
If your school had a math team and you were on it, would be surprised if you didn't hear of it
You may not have heard of the IMO because no one in school district, possibly even state got in. It is extremely selective (like 20 students in the entire country)
Re: OpenAI claims gold-medal performance at IMO 2025
#388Earlier quoted context omitted.
>> Why is that less exciting? A machine competing in an unconstrained natural language difficult math contest and coming out on top by any means is breath taking science fiction a few years ago - now it’s not exciting? Half the internet is convinced that LLMs are a big data cheating machine and if they're right then, yes, boldly cheating where nobody has cheated before is not that exciting.
I don't get it, how do you "big data cheat" an AI into solving previously unencountered problems? Wouldn't that just be engineering?
I haven't read the IMO problems, but knowing how math Olympiad problems work, they're probably not really "unencountered".
People aren't inventing these problems ex nihilo, there's a rulebook somewhere out there to make life easier for contest organizers.
People aren't doing these contests for money, they are doing them for honor, so there is little incentive to cheat. With big business LLM vendors it's a different situation entirely.
Re: OpenAI claims gold-medal performance at IMO 2025
#389[dead]
it would raise more concerns that corps leaked questions/answers to training data and finetuned specialized models during this time.
Re: OpenAI claims gold-medal performance at IMO 2025
#390Earlier quoted context omitted.
> I think we’re close to AI substantially contributing to scientific discovery. The new "Full Self-Driving next year"?
I know it’s a meme but there actually are fully self driving cars, they make thousands of trips every day in a couple US cities.
FWIW, when you get this reductive with your criterion there were technically self-driving cars in 2008 too.