Earlier quoted context omitted.
I am a professor in a math department (I teach statistics but there is a good complement of actual math PhDs) and there are only about 10% who care about these types of problems and definitely less than half who could get gold on an IMO test even if they didn’t care. They are all outstanding mathematicians, but the IMO type questions are not something that mathematicians can universally solve without preparation. The…
> They are all outstanding mathematicians, but the IMO type questions are not something that mathematicians can universally solve without preparation. So IMO is basically the leetcode of Mathematics.
OpenAI claims gold-medal performance at IMO 2025
501–510 of 737 posts
Re: OpenAI claims gold-medal performance at IMO 2025
#502Earlier quoted context omitted.
[flagged]
I am a professor in a math department (I teach statistics but there is a good complement of actual math PhDs) and there are only about 10% who care about these types of problems and definitely less than half who could get gold on an IMO test even if they didn’t care. They are all outstanding mathematicians, but the IMO type questions are not something that mathematicians can universally solve without preparation. The…
Re: OpenAI claims gold-medal performance at IMO 2025
#503Earlier quoted context omitted.
You should basically assume they are pulled from thin air. (Or more precisely, from the brain and world model of the people making the prediction.) The point of giving such estimates is mostly an exercise in getting better at understanding the world, and a way to keep yourself honest by making predictions in advance. If someone else consistently gives higher probabilities to events that ended up happening than you di…
Is there some database where you can see predictions of different people and the results? Or are we supposed to rely on them keeping track and keeping themselves honest? Because that is not something humans do generally, and I have no reason to trust any of these 'rationalists'. This sounds like a circular argument. You started explaining why them giving percentage predictions should make them more trustworthy, but w…
People's bets are publicly viewable. The website is very popular with these "rationality-ists" you refer to.
I wasn't in fact arguing that giving a prediction should make people more trustworthy, please explain how you got that from my comment? I said that the main benefit to making such predictions is as practice for the predictor themselves. If there's a benefit for readers, it is just that they could come along and say "eh, I think the chance is higher than that". Then they also get practice and can compare how they did when the outcome is known.
Re: OpenAI claims gold-medal performance at IMO 2025
#504Pre-registering a prediction: When (not if) AI does make a major scientific discovery, we'll hear "well it's not really thinking, it just processed all human knowledge and found patterns we missed - that's basically cheating!"
Less that AI is cheating and more that we basically found a way to take the thousand monkeys with infinite time scenario and condense that into a reasonable(?) amount of time and with some decent starting instructions. The AI wouldn't have done any of the heavy lifting of the discovery, it just iterated on the work of past researchers at speeds beyond human.
IE, they...
- Start with the context window of prior researchers.
- Set a goal or research direction.
- Engage in chain of thought with occasional reality-testing.
- Generate an output artifact, reviewable by those with appropriate expertise, to allow consensus reality to accept or reject their work.
Re: OpenAI claims gold-medal performance at IMO 2025
#505Earlier quoted context omitted.
I don't get it, how do you "big data cheat" an AI into solving previously unencountered problems? Wouldn't that just be engineering?
I mean, solutions for the 2025 IMO problems are already available on the internet. How can we be sure these are “unencountered” problems?
Re: OpenAI claims gold-medal performance at IMO 2025
#506From Noam Brown https://x.com/polynoamial/status/1946478258968531288 "When you work at a frontier lab, you usually know where frontier capabilities are months before anyone else. But this result is brand new, using recently developed techniques. It was a surprise even to many researchers at OpenAI. Today, everyone gets to see where the frontier is." and "This was a small team effort led by @alexwei_ . He took a resea…
Re: OpenAI claims gold-medal performance at IMO 2025
#507Earlier quoted context omitted.
The key bit here is whether the LLM doing the cherry picking had knowledge of the solution. If it didn't, this is a meaningful result. That's why I'd like more info, but I fear OpenAI is going to try to keep things under wraps.
> If it didn't We kind of have to assume it didn't right? Otherwise bragging about the results makes zero sense and would be outright misleading.
Re: OpenAI claims gold-medal performance at IMO 2025
#508Interesting that the proofs seem to use a limited vocabulary: https://github.com/aw31/openai-imo-2025-proofs/blob/main/pro... Why waste time say lot word when few word do trick :) Also worth pointing out that Alex Wei is himself a gold medalist at IOI.
Are you saying "see the world?" or "seaworld"?
Re: OpenAI claims gold-medal performance at IMO 2025
#509Some previous predictions: In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025. He thought there was an 8% chance of this happening. Eliezer Yudkowsky said "at least 16%". Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...
While I usually enjoy seeing these discussions, I think they are really pushing the usefulness of bayesian statistics. If one dude says the chance for an outcome is 8% and another says it's 16% and the outcome does occur, they were both pretty wrong, even though it might seem like the one who guessed a few % higher might have had a better belief system. Now if one of them had said 90% while the other said 8% or 16%,…
Re: OpenAI claims gold-medal performance at IMO 2025
#510Earlier quoted context omitted.
Implying results are fraudulent is completely fair when it is a fraud. The previous time they had claims about solving all of the math right there and right then, they were caught owning the company that makes that independent test , and could neither admit nor deny training on closed test set.
Just to quickly clarify: - OpenAI doesn't own Epoch AI (though they did commission Epoch to make the eval) - OpenAI denied training on the test set (and further denied training on FrontierMath-derived data, training on data targeting FrontierMath specifically, or using the eval to pick a model checkpoint; in fact, they only downloaded the FrontierMath data after their o3 training set was frozen and they didn't look a…
you just brought several corp statements which are not grounded into any evidence, and could be not true, so you didn't say that much so far.
> Totally fine not to take every company's word at face value, but imo this would be a weird conspiracy for OpenAI, with very high costs on reputation and morale.
prize is XXB of investments and XXXB of valuation, so nothing weird in such conspiracy.