Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

731–737 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#731
post #92

In the RLHF sphere you could tell some AI company/companies were targeting this because of how many IMO RLHF’ers they were hiring specifically. I don’t think it’s really easy to say how much “progress” this is given that.

They were hiring IMO winners because IMO winners tend to be good at working on AI, not because they had the people specifically to make the AI better at math.

Uh no. I’m a math RLHF’er. When I get hired, I work on math/logic up to masters level because that’s my qualifications. Masters and PHD work on masters and PHD level. And IMO work on IMO math.

Every skill and skill level is specifically assigned and hired in the RLHF world.

Sometime the skill levels are fuzzier, but that’s usually very temporary.

And as been said already, IMO is a specific skill that even PHD math holders aren’t universally trained for.

Re: OpenAI claims gold-medal performance at IMO 2025

#732

Earlier quoted context omitted.

Fair points, but the reason everyone is amazed is that five years ago this was entirely impossible for computers irrespective of the competition format or rules. It’s as-if we had learned whale song, and then within two years a whale had won a Nobel prize for their research in high pressure aquatic environments. You’d similarly get naysayers debating the finer points of what special advantage whales may have in that…

And it’s very impressive that whales can write papers. A computer system that can perform these tasks that were unthinkably complex a few years ago is quite impressive. That is a big win, and it can be celebrated. They don’t need to be celebrated as a “gold medalist” if they didn’t perform according to the same criteria as a gold-medalist.

Win for whom?

Re: OpenAI claims gold-medal performance at IMO 2025

#733

Earlier quoted context omitted.

Sure, but nobody is using their IMO score to prove they are superintelligent and pulling it off in wider groups.

I’ve seen IMO rank used to justify more than one $100m+ seed round.

Someone was burned by Cognition?

Re: OpenAI claims gold-medal performance at IMO 2025

#734
post #215

Earlier quoted context omitted.

With regard to AI & LLMs Twitter/x is actually the only place with all of the industry people discussing. There are a bunch of great accounts to follow that are only really posting content to x. Karpathy, nearcyan, kalomaze, all of the OpenAI researchers including the link this discussion is on, many anthropic researchers. It's such a meme that you see people discuss reading Twitter thread + paper because the thread…

Too hyperbolic for, against, or either way?

It honestly depends on the headline.

I think hn probably has a disproportionate number of haters while Twitter has a disproportionate number of blind believers / hype types.

But both have both.

Not sure how this compares to YouTube (although my guess is the thumbnails + titles are most egregious there for algorithm reasons)

Re: OpenAI claims gold-medal performance at IMO 2025

#735

From Noam Brown https://x.com/polynoamial/status/1946478258968531288 "When you work at a frontier lab, you usually know where frontier capabilities are months before anyone else. But this result is brand new, using recently developed techniques. It was a surprise even to many researchers at OpenAI. Today, everyone gets to see where the frontier is." and "This was a small team effort led by @alexwei_ . He took a resea…

That brand new technique? Training on the test data. /s

Do you have any proof to support this claim?

Re: OpenAI claims gold-medal performance at IMO 2025

#736

Earlier quoted context omitted.

Hence proofs as I've stated.

Go up to Andrew Wiles and say, "Meh, NBD, it was just a proof."

IMO questions and Andrew Wiles solving Fermat's last theorem are two vastly different things. One is far harder than the other and the effort he put in and thinking needed is something very few can do. He also did some other fascinating work that I couldn't hope to understand fully. There is a gulf between FLT and IMO types of proofs.

Re: OpenAI claims gold-medal performance at IMO 2025

#737

Earlier quoted context omitted.

A thought-terminating cliché? Not at all, certainly not when it comes to claims of technological or scientific breakthroughs. After all, that's partly why we have peer review and an emphasis on reproducibility. Until such a claim has been scrutinised by experts or reproduced by the community at large, it remains an unverified claim. >> Unlike seemingly most here on HN, I judge people's trustworthiness individually an…

They don't give a lot of details but they give enough for it to be pretty hard to say the claim is false but unfraudulent. Some researchers got a breakthrough and decided to share right then rather than the months later it would take for a viable product. It happens, researchers are humans after all and i'm generally glad to take a peek at the actual frontier rather than what's behind by many months. You can and it's…

I'm not ignoring it. I'm waiting to see evidence of it. Is that uncharitable?
Post reply on HN