Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

311–320 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#311

Earlier quoted context omitted.

> These models cannot reason Not trying to be a smarty pants here, but what do we mean by "reason"? Just to make the point, I'm using Claude to help me code right now. In between prompts, I read HN. It does things for me such as coding up new features, looking at the compile and runtime responses, and then correcting the code. All while I sit here and write with you on HN. It gives me feedback like "lock free message…

You raise a far point. These criticisms based on "it's merely X" or "it's not really Y" don't hold water when X and Y are poorly defined. The only thing that should matter is the results they get. And I have a hard time understanding why the thing that is supposed to behave in an intelligent way but often just spew nonsense gets 10x budget increases over and over again. This is bad software. It does not do the thing…

You're right, they should never have given more resources and compute to the OpenAI team after the disaster called GPT-2, which only knew how to spew nonsense.

Re: OpenAI claims gold-medal performance at IMO 2025

#312
post #160

Earlier quoted context omitted.

A third party tried this experiment with publicly available models. OpenAI did half as well as Gemini, and none of the models even got bronze. https://matharena.ai/imo/

I feel you're misunderstanding something. That's not "this exact experiment". Matharena is testing publicly available models against the IMO problem set. OpenAI was announcing the results of a new, unpublished model, on that problems set. It is totally fair to discount OpenAI's statement until we have way more details about their setup, and maybe even until there is some level of public access to the model. But you'r…

Implying results are fraudulent is completely fair when it is a fraud.

The previous time they had claims about solving all of the math right there and right then, they were caught owning the company that makes that independent test, and could neither admit nor deny training on closed test set.

Re: OpenAI claims gold-medal performance at IMO 2025

#313

Noam Brown: > this isn’t an IMO-specific model. It’s a reasoning LLM that incorporates new experimental general-purpose techniques. > it’s also more efficient [than o1 or o3] with its thinking. And there’s a lot of room to push the test-time compute and efficiency further. > As fast as recent AI progress has been, I fully expect the trend to continue. Importantly, I think we’re close to AI substantially contributing…

I assume there was tool use in the fine tuning?

There wasn’t in the CoT for these problems.

Re: OpenAI claims gold-medal performance at IMO 2025

#315

Noam Brown: > this isn’t an IMO-specific model. It’s a reasoning LLM that incorporates new experimental general-purpose techniques. > it’s also more efficient [than o1 or o3] with its thinking. And there’s a lot of room to push the test-time compute and efficiency further. > As fast as recent AI progress has been, I fully expect the trend to continue. Importantly, I think we’re close to AI substantially contributing…

What's the clear path to improved efficiency now that we've reached peak data?

Re: OpenAI claims gold-medal performance at IMO 2025

#316

These are high school level only in the sense of assumed background knowledge, they are extremely difficult. Professional mathematicians would not get this level of performance, unless they have a background in IMO themselves. This doesn’t mean that the model is better than them in math, just that mathematicians specialize in extending the frontier of math. The answers are not in the training data. This is not a mode…

[flagged]

This is probably the right time to bring up this classic:

"Did you win the Putnam?"

https://news.ycombinator.com/item?id=35079

Re: OpenAI claims gold-medal performance at IMO 2025

#317
post #9

Some previous predictions: In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025. He thought there was an 8% chance of this happening. Eliezer Yudkowsky said "at least 16%". Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...

While I usually enjoy seeing these discussions, I think they are really pushing the usefulness of bayesian statistics. If one dude says the chance for an outcome is 8% and another says it's 16% and the outcome does occur, they were both pretty wrong, even though it might seem like the one who guessed a few % higher might have had a better belief system. Now if one of them had said 90% while the other said 8% or 16%,…

The correctness of 8%, 16%, and 90% are all equally unknown since we only have one timeline, no?

Re: OpenAI claims gold-medal performance at IMO 2025

#318

These are high school level only in the sense of assumed background knowledge, they are extremely difficult. Professional mathematicians would not get this level of performance, unless they have a background in IMO themselves. This doesn’t mean that the model is better than them in math, just that mathematicians specialize in extending the frontier of math. The answers are not in the training data. This is not a mode…

> The answers are not in the training data.

> This is not a model specialized to IMO problems.

Any proof?

Re: OpenAI claims gold-medal performance at IMO 2025

#319

Earlier quoted context omitted.

How is a claim , "clear evidence" to anything?

Most evidence you have about the world is claims from other people, not direct experiment. There seems to be a thought-terminating cliche here on HN, dismissing any claim from employees of large tech companies. Unlike seemingly most here on HN, I judge people's trustworthiness individually and not solely by the organization they belong to. Noam Brown is a well known researcher in the field and I see no reason to doub…

> dismissing any claim from employees of large tech companies

Me: I have a way to turn lead into gold.

You: Show me!!!

Me: NO (and then spends the rest of my life in poverty).

Cold Fusion (physics not the programing language) is the best example of why you "Show your work". This is the Valley we're talking about. It's the thudnderdome of technology and companies. If you have a meaningful breakthrough you don't talk about it you drop it on the public and flex.

Re: OpenAI claims gold-medal performance at IMO 2025

#320
post #267

Earlier quoted context omitted.

Making an account just to point out how these comments are far more exhausting, because they don't engage with the subject matter. They are just agreeing with a headline and saying, "See?" You say, "explaining away the increasing performance" as though that was a good faith representation of arguments made against LLMs, or even this specific article. Questionong the self-congragulatory nature of these businesses is p…

But don't you think this might be a case where there is both self-congragulation and actual progress?

The level of proof for the latter is much higher, and IMO, OpenAI hasn't met the bar yet.

Something really funky is going on with newer AI models and benchmarks, versus how they perform subjectively when I use them for my use-cases. I say this across the board[1], not just regarding IpenAI. I don't know if frontier labs have run into Goodheart's law viz benchmarks, or if my use-cases that are atypical.

1. I first noticed this with Claud 3.5 vs Claud 3.7

Post reply on HN