Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

501–510 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#501
post #367

Earlier quoted context omitted.

I am a professor in a math department (I teach statistics but there is a good complement of actual math PhDs) and there are only about 10% who care about these types of problems and definitely less than half who could get gold on an IMO test even if they didn’t care. They are all outstanding mathematicians, but the IMO type questions are not something that mathematicians can universally solve without preparation. The…

> They are all outstanding mathematicians, but the IMO type questions are not something that mathematicians can universally solve without preparation. So IMO is basically the leetcode of Mathematics.

Yeah, no - quite a chunk of IMO problems are planar and 3d geometry, and you don't really do that at university level (exception: specializing in high school maths didactics)

Re: OpenAI claims gold-medal performance at IMO 2025

#502

Earlier quoted context omitted.

[flagged]

I am a professor in a math department (I teach statistics but there is a good complement of actual math PhDs) and there are only about 10% who care about these types of problems and definitely less than half who could get gold on an IMO test even if they didn’t care. They are all outstanding mathematicians, but the IMO type questions are not something that mathematicians can universally solve without preparation. The…

So IMO questions are to math what Leetcode is to programming?

Re: OpenAI claims gold-medal performance at IMO 2025

#503

Earlier quoted context omitted.

You should basically assume they are pulled from thin air. (Or more precisely, from the brain and world model of the people making the prediction.) The point of giving such estimates is mostly an exercise in getting better at understanding the world, and a way to keep yourself honest by making predictions in advance. If someone else consistently gives higher probabilities to events that ended up happening than you di…

Is there some database where you can see predictions of different people and the results? Or are we supposed to rely on them keeping track and keeping themselves honest? Because that is not something humans do generally, and I have no reason to trust any of these 'rationalists'. This sounds like a circular argument. You started explaining why them giving percentage predictions should make them more trustworthy, but w…

Yes, there is: https://manifold.markets/

People's bets are publicly viewable. The website is very popular with these "rationality-ists" you refer to.

I wasn't in fact arguing that giving a prediction should make people more trustworthy, please explain how you got that from my comment? I said that the main benefit to making such predictions is as practice for the predictor themselves. If there's a benefit for readers, it is just that they could come along and say "eh, I think the chance is higher than that". Then they also get practice and can compare how they did when the outcome is known.

Re: OpenAI claims gold-medal performance at IMO 2025

#504
post #500
post #490

Pre-registering a prediction: When (not if) AI does make a major scientific discovery, we'll hear "well it's not really thinking, it just processed all human knowledge and found patterns we missed - that's basically cheating!"

Less that AI is cheating and more that we basically found a way to take the thousand monkeys with infinite time scenario and condense that into a reasonable(?) amount of time and with some decent starting instructions. The AI wouldn't have done any of the heavy lifting of the discovery, it just iterated on the work of past researchers at speeds beyond human.

Honest question - how is that not true of those past researchers?

IE, they...

- Start with the context window of prior researchers.

- Set a goal or research direction.

- Engage in chain of thought with occasional reality-testing.

- Generate an output artifact, reviewable by those with appropriate expertise, to allow consensus reality to accept or reject their work.

Re: OpenAI claims gold-medal performance at IMO 2025

#505

Earlier quoted context omitted.

I don't get it, how do you "big data cheat" an AI into solving previously unencountered problems? Wouldn't that just be engineering?

I mean, solutions for the 2025 IMO problems are already available on the internet. How can we be sure these are “unencountered” problems?

They probably have an archived data set from before then that they trained on.

Re: OpenAI claims gold-medal performance at IMO 2025

#506

From Noam Brown https://x.com/polynoamial/status/1946478258968531288 "When you work at a frontier lab, you usually know where frontier capabilities are months before anyone else. But this result is brand new, using recently developed techniques. It was a surprise even to many researchers at OpenAI. Today, everyone gets to see where the frontier is." and "This was a small team effort led by @alexwei_ . He took a resea…

[deleted]

Re: OpenAI claims gold-medal performance at IMO 2025

#507
post #86

Earlier quoted context omitted.

The key bit here is whether the LLM doing the cherry picking had knowledge of the solution. If it didn't, this is a meaningful result. That's why I'd like more info, but I fear OpenAI is going to try to keep things under wraps.

> If it didn't We kind of have to assume it didn't right? Otherwise bragging about the results makes zero sense and would be outright misleading.

Corporations mislead to make money all the damn time.

Re: OpenAI claims gold-medal performance at IMO 2025

#508

Interesting that the proofs seem to use a limited vocabulary: https://github.com/aw31/openai-imo-2025-proofs/blob/main/pro... Why waste time say lot word when few word do trick :) Also worth pointing out that Alex Wei is himself a gold medalist at IOI.

Are you saying "see the world?" or "seaworld"?

[deleted]

Re: OpenAI claims gold-medal performance at IMO 2025

#509
post #9

Some previous predictions: In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025. He thought there was an 8% chance of this happening. Eliezer Yudkowsky said "at least 16%". Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...

While I usually enjoy seeing these discussions, I think they are really pushing the usefulness of bayesian statistics. If one dude says the chance for an outcome is 8% and another says it's 16% and the outcome does occur, they were both pretty wrong, even though it might seem like the one who guessed a few % higher might have had a better belief system. Now if one of them had said 90% while the other said 8% or 16%,…

The whole point is to make many such predictions and experience many outcomes. The goal is for your 70% predictions to be correct 70% of the time. We all have a gap between how confident we are and how often we're correct. Calibration, which can be measured by making many predictions, is about reducing that gap.

Re: OpenAI claims gold-medal performance at IMO 2025

#510

Earlier quoted context omitted.

Implying results are fraudulent is completely fair when it is a fraud. The previous time they had claims about solving all of the math right there and right then, they were caught owning the company that makes that independent test , and could neither admit nor deny training on closed test set.

Just to quickly clarify: - OpenAI doesn't own Epoch AI (though they did commission Epoch to make the eval) - OpenAI denied training on the test set (and further denied training on FrontierMath-derived data, training on data targeting FrontierMath specifically, or using the eval to pick a model checkpoint; in fact, they only downloaded the FrontierMath data after their o3 training set was frozen and they didn't look a…

> if that's you feel there's probably not much I can say to change your mind

you just brought several corp statements which are not grounded into any evidence, and could be not true, so you didn't say that much so far.

> Totally fine not to take every company's word at face value, but imo this would be a weird conspiracy for OpenAI, with very high costs on reputation and morale.

prize is XXB of investments and XXXB of valuation, so nothing weird in such conspiracy.

Post reply on HN