Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

691–700 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#691

Terence Tao on the matter - https://imgur.com/a/terence-tao-on-supposed-gold-imo-sMKP0bm

> there will be a proposal at some point to actually have an AI math Olympiad where at the same time as the human contestants get the actual Olympiad problems, AI’s will also be given the same problems, the same time period and the outputs will have to be graded by the same judges, which means that it’ll have be written in natural language rather than formal language.[1]

Last month, Tao himself said that we can compare humans and AIs at IMO. He even said such AI didn't exist yet and AIs won't beat IMO in 2025. And now that AIs can compete with humans at IMO under the same conditions that Tao mentioned, suddenly it becomes an apples-to-oranges comparison?

[1] https://lexfridman.com/terence-tao-transcript/

Re: OpenAI claims gold-medal performance at IMO 2025

#692
post #678

Earlier quoted context omitted.

To the nearest 5%, for percentages in that middle range. It is not just 16% -> 15%. But also 46% -> 45%.

Yes so this confirms my point rather than refuting it...

It seems that you reversed your point then. You said before:

Even before I read your comment I thought that 5% precision is useful but 1% precision is a silly turn-off, unless that 1% is near the 0% or 100% boundary.

However what I am saying is that there is real data, involving real predictions, by real people, that demonstrates that there is a measurable statistical loss of accuracy in their predictions if you round off those percentages.

This doesn't mean that any individual prediction is accurate to that percent. But it happens often enough that the last percent really does contain real value.

Re: OpenAI claims gold-medal performance at IMO 2025

#693
post #686

Earlier quoted context omitted.

It's "not interesting" because no novel insight has to be used in order to solve this. It's immediately obvious how to solve it, just follow the textbook procedure. This is distinct both from other typical IMO problems that I've seen and from research mathematics which usually do require some amount of creativity. > exp(i\pi)+1=0 If your definition of "exp(i*theta)" is literally "rotation of the number 1 by theta deg…

Those definitions of exp are all immediately obvious and nearly the definition of textbook, every university calculus course covers them. That is the issue with defining interesting as novel - nothing generally known is novel any more. And they don't require any special maths - sum_{i=0}^\infty z^n/n! is literally just multiplication and addition. The long and short of it is it just isn't possible to tell someone tha…

> those definitions of exp are all immediately obvious

no

Re: OpenAI claims gold-medal performance at IMO 2025

#694
post #353

Earlier quoted context omitted.

That's why you have to let these people make predictions about many things. Than you can weigh the 8, 16, and 90 pct and see who is talking out of their ass.

That's just the frequentist approach. But we're talking about bayesian statistics here.

I admit I dont know Bayesian, but isn't the only way to check if the future teller is lucky or not to have them predict many things? If he predicts 10 to happen with a 10% chance, and one of them happens, he's good. If he predicts 10 to happen with a 90% chance and 9 happen, same. How is this different with Bayesian?

Re: OpenAI claims gold-medal performance at IMO 2025

#696

Earlier quoted context omitted.

In the IMO, the idea is that the first day you get P1, P2 and P3, and the second day you get P4, P5 and P6. Usually, ordered by difficulty, they are P1, P4, P2, P5, P3, P6. So, usually P1 is "easy" and P6 is very hard. At least that is the intended order, but sometime reality disagree. Edit: Fixed P4 -> P3. Thanks.

In this case P6 was unusually hard and P3 was unusually easy https://sugaku.net/content/imo-2025-problems/

Yikes. 30 years ago I would eat this stuff up and I was the lead dev on 3D engines.

Now I can't even make heads-or-tails of what P6 is even asking (^▽^)

Re: OpenAI claims gold-medal performance at IMO 2025

#697

Noam Brown: > this isn’t an IMO-specific model. It’s a reasoning LLM that incorporates new experimental general-purpose techniques. > it’s also more efficient [than o1 or o3] with its thinking. And there’s a lot of room to push the test-time compute and efficiency further. > As fast as recent AI progress has been, I fully expect the trend to continue. Importantly, I think we’re close to AI substantially contributing…

What's the clear path to improved efficiency now that we've reached peak data?

We're so far from peak data that we've barely even scratched the surface, IMO.

Re: OpenAI claims gold-medal performance at IMO 2025

#698
post #635

Earlier quoted context omitted.

I think you misundertand me, I'm making some pie in the sky statement about AI being able to discover the laws of nature in an afternoon. I'm just making the observation that if you know the basic equiations, and enough math (which is about multivariate calc), you can derive every single formula in your Physics textbook (and most undergrads do as part of their education). Since smart people can derive a lot of knowle…

There's no evidence this model works like that. The "axioms" for counting the number of r's in a word are magnitudes simpler than classical physic's, and yet it took a few years to get that right. It's always been context, not derivation of logic.

First, false equivalence. The 'strawberry' problem was because LLMs operate not on text directly, but on embedding vectors, which made it hard for it to manipulate the syntax of language directly. This does not prevent it from properly doing math proofs.

Second, we know nothing about these models or how they work and trained, and indeed, if they can do these things or not. But a smart human could (by smart I mean someone who gets good grades at engineering school effortlessly, not Albert Einstein)

Re: OpenAI claims gold-medal performance at IMO 2025

#699

Earlier quoted context omitted.

Imo competitive math (or programming) is about knowing some tricks and then trying to find a combination of them that works for a given task. The number of tricks and depth required is much less than in go or chess. I don't think it's very creative endeavor in comparison to chess/go. The searching required is less as well. There is a challenge processing natural language and producing solutions in it though. Creativi…

I'd disagree with this take. Math olympiads are some of the most intellectually creative activities I've ever done that fit within a one day time limit. Chess and go don't even come close--I am not a strong player, but I've studied both games for hundreds of hours. (My hot take is that chess is not even very creative at all, that's why classical AI techniques produced super human results many years ago.) There is no…

What you are missing about chess and go is that those games are not about finding one true solution. They are very psychological games (at human level) and are about finding moves that are difficult to handle for the opponent. You try to understand how your opponent thinks and what is going to be unpleasant for them. This gives a lot of scope for creative and psychological warfare.

In competitive math (or programming) there is one correct solution and no opponent. It's just not possible for it to be very creative endeavor if those solutions can be found in very limited time.

>>(I've heard a contestant say that IMO prep was memorizing a lot of template solutions, but he was such a genius among geniuses that I think his opinion is irrelevant to the rest of humanity!)

So you have not only chosen to ignore the view of someone who is very good at it but also assumed that even though the best preparation for them is to memorize a lot of solutions it must be about creativity for people who are not geniuses like this guy? How does it make sense at all?

Re: OpenAI claims gold-medal performance at IMO 2025

#700

Earlier quoted context omitted.

I don't fault you for maintaining a healthy scepticism, but per the President of the IMO: "It is very exciting to see progress in the mathematical capabilities of AI models, but we would like to be clear that the IMO cannot validate the methods, including the amount of compute used or whether there was any human involvement, or whether the results can be reproduced. What we can say is that correct mathematical proofs…

Is b) really that unlikely?

Not really. This whole thing looks like a deliberately planned PR campaign, similar to the o3 demo. OpenAI has enough talented mathematicians. They had enough time to just solve the problems themselves. Alternatively, some participants leaking the questions for a reward isn't very unlikely either, and I definitely wouldn't put it past OpenAI to try something like that. Afterwards, they could secretly give hints or tool access to the model, or simply forge the answers, or keep rerunning the model until it gave out the correct answer. We know from FrontierMath and ARC-AGI that OpenAI can't be trusted when it comes to benchmarks.
Post reply on HN