Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

561–570 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#561

Earlier quoted context omitted.

To me, this is a tell of human-involvement in the model solution. There is no reason why machines would do badly on exactly the problem which humans do badly as well - without humans prodding the machine towards a solution. Also, there is no reason why machines could not produce a partial or wrong answer to problem 6 which seems like survivor bias to me. ie, that only correct solutions were cherrypicked.

Maybe it’s a hint that our current training techniques can create models comparable to the best humans in a given subject, but that’s the limit.

We've hit the limit of 'our current training techniques'? This result literally used newly developed techniques that surprised researchers at OpenAI.

Noam Brown: 'This result is brand new, using recently developed techniques. It was a surprise even to many researchers at OpenAI.'

So your thesis is that these new techniques - which just produced unexpected breakthroughs - represent some kind of ceiling? That's an impressive level of confidence about the limits of methods we apparently just invented in a field which seems to, if anything, be accelerating.

Re: OpenAI claims gold-medal performance at IMO 2025

#562
post #210

The cynicism/denial on HN about AI is exhausting. Half the comments are some weird form of explaining away the ever increasing performance of these models I've been reading this website for probably 15 years, its never been this bad. many threads are completely unreadable, all the actual educated takes are on X, its almost like there was a talent drain

> I've been reading this website for probably 15 years, its never been this bad. People here were pretty skeptical about AlexNet, when it won the ImageNet challenge 13 years ago. https://news.ycombinator.com/item?id=4611830

Ouch that thread makes me quite sad at the state of discourse on HN today. It's a lot better than this thread.

Re: OpenAI claims gold-medal performance at IMO 2025

#563
post #490

Pre-registering a prediction: When (not if) AI does make a major scientific discovery, we'll hear "well it's not really thinking, it just processed all human knowledge and found patterns we missed - that's basically cheating!"

Turns out goalposts are the world’s most easily moved objects. We should start building spacecraft out of them.

I saw the phrase "goalposts aren't just moving, they're doing parkour" recently and I do love that image. It does seem to capture the state of things quite well.

Re: OpenAI claims gold-medal performance at IMO 2025

#564

Earlier quoted context omitted.

Much more efficient for us to all speak the same language. Trying to create fragmentation is inefficient.

You should take that up with the IMO then, or all of European Union. They provide services in ~two dozen languages.

Sure, but why worsen the situation by using more languages?

Re: OpenAI claims gold-medal performance at IMO 2025

#565

Earlier quoted context omitted.

I think the biggest hint that the models aren't reasoning is that they can't explain their reasoning. Researchers have shown for explained that how a model solves a simple math problem and how it claims to have solved it after the fact have no real correlation. In other words there was only the appearance of reasoning.

People can't explain their reasoning either. People do a parallel construction of logical arguments for a conclusion they already reached intuitively in a way they have no clue how it happened. "The idea just popped into my head while showering" to our credit, if this post-hoc rationalization fails we are able to change our opinion to some degree.

Interestingly people have to be trained in logic and identifying fallacies because logic is not a native capability of our mind. We aren’t even that good at it once trained and many humans (don’t forget a 100 IQ is median) can not be trained.

Reasoning appears to actually be more accurately described as “awareness,” or some process that exists along side thought where agency and subconscious processes occur. It’s by construction unobservable by our conscious mind, which is why we have so much trouble explaining it. It’s not intuition - it’s awareness.

Re: OpenAI claims gold-medal performance at IMO 2025

#566

Earlier quoted context omitted.

Off topic, but am I the only one getting triggered every time I see a rationalist quantify their prediction of the future with single digit accuracy? It's like their magic way of trying to get everyone to forget that they reached their conclusion in completely hand-wavy way, just like every other human being. But instead of saying "low confidence" or "high confidence" like the rest of us normies, they will tell you t…

Interestingly, this is actually a question that's been looked at empirically! Take a look at this paper: https://scholar.harvard.edu/files/rzeckhauser/files/value_of... They took high-precision forecasts from a forecasting tournament and rounded them to coarser buckets (nearest 5%, nearest 10%, nearest 33%), to see if the precision was actually conveying any real information. What they found is that if you rounded th…

[deleted]

Re: OpenAI claims gold-medal performance at IMO 2025

#567

These are high school level only in the sense of assumed background knowledge, they are extremely difficult. Professional mathematicians would not get this level of performance, unless they have a background in IMO themselves. This doesn’t mean that the model is better than them in math, just that mathematicians specialize in extending the frontier of math. The answers are not in the training data. This is not a mode…

It almost certainly is specialized to IMO problems, look at the way it is answering the questions: https://xcancel.com/alexwei_/status/1946477742855532918 E.g here: https://pbs.twimg.com/media/GwLtrPeWIAUMDYI.png?name=orig Frankly it looks to me like it's using an AlphaProof style system, going between natural language and Lean/etc. Of course OpenAI will not tell us any of this.

Since this looks like geometric proof, I wonder if the AI operates only on logical/mathematical statements or it actually somehow 'visualizes' the proof like a human would while solving.

Re: OpenAI claims gold-medal performance at IMO 2025

#568
post #490

Pre-registering a prediction: When (not if) AI does make a major scientific discovery, we'll hear "well it's not really thinking, it just processed all human knowledge and found patterns we missed - that's basically cheating!"

People get very fragile when AI is better at something than them (excluding speed/scale of operations, where computers have an obvious edge)

Re: OpenAI claims gold-medal performance at IMO 2025

#569

These are high school level only in the sense of assumed background knowledge, they are extremely difficult. Professional mathematicians would not get this level of performance, unless they have a background in IMO themselves. This doesn’t mean that the model is better than them in math, just that mathematicians specialize in extending the frontier of math. The answers are not in the training data. This is not a mode…

> The answers are not in the training data. > This is not a model specialized to IMO problems. Any proof?

No, and they're lying on the most important claim: that this is not a model specialized to IMO problems.

From the thread:

> just to be clear: the IMO gold LLM is an experimental research model.

The thread tried to muddy the narrative by saying the methodology can generalize, but no one is claiming the actual model is a generalized model.

There'd be a massively different conversation needed if a generalized model that could become the next iteration of ChatGPT had achieved this level of performance.

Post reply on HN