Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

631–640 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#631
post #373

Earlier quoted context omitted.

Can you provide any kind of source? Very curious about this!

https://work.mercor.com/jobs/list_AAABljpKHPMmFMXrg2VM0qz4 https://benture.io/job/international-math-olympiad-participa... https://job-boards.greenhouse.io/xai/jobs/4538773007 And Outlier/Scale, which was bought by Meta (via Scale), had many IMO-required Math AI trainer jobs on LinkedIn. I can't find those historical ones though. I'm just one piece in the cog and this is an anecdote, but there was a huge upswing in I…

I would fully expect every IMO participant grinds IMO problems for months before the competition.

I don't know why people hold training a model on like material as a negation of it's ability.

Re: OpenAI claims gold-medal performance at IMO 2025

#632
post #490

Pre-registering a prediction: When (not if) AI does make a major scientific discovery, we'll hear "well it's not really thinking, it just processed all human knowledge and found patterns we missed - that's basically cheating!"

It's stochastic parrots all the ways down:

https://ai.vixra.org/pdf/2506.0065v1.pdf

A satirical paper, but it's hilariously brilliant.

Re: OpenAI claims gold-medal performance at IMO 2025

#633

Earlier quoted context omitted.

Google’s AlphaProof, which got a silver last year, has been using a neural symbolic approach. This gold from OpenAI was pure LLM. We’ll have to see what Google announces, but the LLM approach is interesting because it will likely generalize to all kinds of reasoning problems, not just mathematical proofs.

OpenAI’s systems haven’t been pure language models since the o models though, right? Their RL approach may very well still generalize, but it’s not just a big pre-trained model that is one-shotting these problems. The key difference is that they claim to have not used any verifiers.

What do you mean by “pure language model”? The reasoning step is still just the LLM spitting out tokens and this was confirmed by Deepseek replicating the o models. There’s not also a proof verifier or something similar running alongside it according to the openai researchers.

If you mean pure as in there’s not additional training beyond the pretraining, I don’t think any model has been pure since gpt-3.5.

Re: OpenAI claims gold-medal performance at IMO 2025

#634

Earlier quoted context omitted.

Why is that less exciting? A machine competing in an unconstrained natural language difficult math contest and coming out on top by any means is breath taking science fiction a few years ago - now it’s not exciting? Regardless of the tools for verification or even solvers - why is the goal post moving so fast? There is no bonus for “purity of essence” and using only neural networks. We live in an era where it’s hard…

>Why is that less exciting? Because if I have to throw 10000 rocks to get one in the bucket, I am not as good/useful of a rock-into-bucket-thrower as someone who gets it in one shot. People would probably not be as excited about the prospect of employing me to throw rocks for them.

It’s exciting because nearly all humans have 0% chance of throwing the rock into the bucket, and most people believed a rock-into-bucket-thrower machine is impossible. So even an inefficient rock-into-bucket-thrower is impressive.

But the bar has been getting raised very rapidly. What was impressive six months ago is awful and unexciting today.

Re: OpenAI claims gold-medal performance at IMO 2025

#635
post #604

Earlier quoted context omitted.

If you look at the history of physics I don't think it really worked like that. It took about three centuries from Newton to Maxwell because it's hard to just deduce everything from basic principles.

I think you misundertand me, I'm making some pie in the sky statement about AI being able to discover the laws of nature in an afternoon. I'm just making the observation that if you know the basic equiations, and enough math (which is about multivariate calc), you can derive every single formula in your Physics textbook (and most undergrads do as part of their education). Since smart people can derive a lot of knowle…

There's no evidence this model works like that. The "axioms" for counting the number of r's in a word are magnitudes simpler than classical physic's, and yet it took a few years to get that right. It's always been context, not derivation of logic.

Re: OpenAI claims gold-medal performance at IMO 2025

#636
post #85

I encourage anyone who thinks these are easy high-school problems to try to solve some. They're published (including this year's) at https://www.imo-official.org/problems.aspx . They make my head spin.

How do those compare to leetcode hard problems?

Re: OpenAI claims gold-medal performance at IMO 2025

#637
post #498

Earlier quoted context omitted.

Having spent tens of thousands of hours contributing to scientific discovery by reading dense papers for a single piece of information, reverse engineering code written by biologists, and tweaking graphics to meet journal requirements… I can say with certainty it’s already contributing by allowing scientists to spend time on science versus spending an afternoon figuring out which undocumented argument in a R package…

That is not what they mean by contributing to scientific discovery.

Perhaps not, but my point stands from personal experience and knowing what’s going on in labs right now that AI is greatly contributing to research even if it’s not doing the parts that most people think of when they think science. A sufficiently advanced AI in the near term isn’t going to start churning out novel hypotheses and being able to collect non-online data without first being able to secure funding to hire grad students or whatever robots can replace those.

Re: OpenAI claims gold-medal performance at IMO 2025

#638
post #85

I encourage anyone who thinks these are easy high-school problems to try to solve some. They're published (including this year's) at https://www.imo-official.org/problems.aspx . They make my head spin.

How do those compare to leetcode hard problems?

Depends on how hard, but the “average hard” leetcode problem is much easier. These will be more like the ACM ICPC level questions, which I’d put at the “hard hard” leetcode level (also this is a collegiate competition rather than high school, but with broader participation).

Re: OpenAI claims gold-medal performance at IMO 2025

#639

Earlier quoted context omitted.

[flagged]

I am a professor in a math department (I teach statistics but there is a good complement of actual math PhDs) and there are only about 10% who care about these types of problems and definitely less than half who could get gold on an IMO test even if they didn’t care. They are all outstanding mathematicians, but the IMO type questions are not something that mathematicians can universally solve without preparation. The…

I see this distinction a lot, but what is the fundamental difference between competition "math" and professional/research math? If people actually knew then they (young students, and their parents) could decide for themselves if they wanted to engage in either kind of study.

Re: OpenAI claims gold-medal performance at IMO 2025

#640

Google also joined IMO, and got gold prize. https://x.com/natolambert/status/1946569475396120653 OAI announced early, probably we will hear announcement from Google soon.

Given the Noam Brown comment ("It was a surprise even to many researchers at OpenAI") it seems extra surprising if multiple labs achieved this result at once.

There's a comment on this twitter thread saying the Google model was using Lean, while IIUC the OpenAI one was pure LLM reasoning (no tools). Anyone have any corroboration?

In a sense it's kinda irrelevant, I care much more about the concrete things AI can achieve, than the how. But at the same time it's very informative to see the limits of specific techniques expand.

Post reply on HN