Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

541–550 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#541
post #205

Earlier quoted context omitted.

Makes sense. Everyone here has their pride and identity tied to their ability to code. HN likes to upvote articles related to IQ because coding correlates with IQ and HNers like to think they are smart. AI is of course a direct attack on the average HNers identity. The response you see is like attacking a Christian on his religion. The pattern of defense is typical. When someone’s identity gets attacked they need to…

[flagged]

> The fact we're now promoting and discussing fucking Twitter threads is absurd.

The ML community is really big on Twitter. I'm honestly quite surprised that you're angry or surprised at this. That means either you're very disconnected from the actual ML community, which is fine of course but then maybe you should hold your opinions a bit less tightly. Alternatively you're ideologically against Twitter which brings me to:

> It's not about pride and identity, you dingus.

Maybe it is? There's a very-online-tech-person identity that I'm familiar with that hates Twitter because they think that Twitter's short post length and other cultural factors on the site contributed to bad discourse quality. I used to buy it, but I've stopped because HN and Reddit are equally filled with terrible comments that generate more heat than light.

FWIW a bunch of ML researchers tried to switch to Bluesky but got so much hate, including death threats, sent at them that they all noped back to Twitter. That's the other identity portion of it that, post Musk there's a set of folks who hate Twitter ideologically and have built an identity around it. Unfortunately this identity also is anti-AI enough that it's willing to act with toxicity toward ML researchers. Tech cynicism and anti-capitalism has some tie-ins with this also.

So IMO there is an identity aspect to this. It might not be the "true hacker" identity that the GP talks about but I do very much think that this pro vs anti AI fight has turned into another culture war axis on HN that has more to do with your identity or tribe than any reasoned arguments.

Re: OpenAI claims gold-medal performance at IMO 2025

#542
post #44

From that thread: "The model solved P1 through P5; it did not produce a solution for P6." It's interesting that it didn't solve the problem that was by far the hardest for humans too. China, the #1 team got only 21/42 points on it. In most other teams nobody solved it.

To me, this is a tell of human-involvement in the model solution. There is no reason why machines would do badly on exactly the problem which humans do badly as well - without humans prodding the machine towards a solution. Also, there is no reason why machines could not produce a partial or wrong answer to problem 6 which seems like survivor bias to me. ie, that only correct solutions were cherrypicked.

Lmao.

You know IMO questions are not all equally difficult, right? They're specifically designed to vary in difficulty. The reason that problem 6 is hard for both humans and LLM is... it's hard! What a surprise.

Re: OpenAI claims gold-medal performance at IMO 2025

#543
post #9

Some previous predictions: In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025. He thought there was an 8% chance of this happening. Eliezer Yudkowsky said "at least 16%". Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...

16% is just a way of saying one in six chances

Or just “twice as likely as the guy who said 8%”.

Re: OpenAI claims gold-medal performance at IMO 2025

#544
post #91
post #9

Some previous predictions: In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025. He thought there was an 8% chance of this happening. Eliezer Yudkowsky said "at least 16%". Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...

Context? Who are these people and what are these numbers and why shouldn't I assume they're pulled from thin air?

>Who are these people

Be glad you don't know anything about them. Seriously.

Re: OpenAI claims gold-medal performance at IMO 2025

#545
post #85

I encourage anyone who thinks these are easy high-school problems to try to solve some. They're published (including this year's) at https://www.imo-official.org/problems.aspx . They make my head spin.

Related — these videos give a sense of how someone might actually go about thinking through and solving these kinds of problems: - A 3Blue1Brown video on a particularly nice and unexpectedly difficult IMO problem (2011 IMO, Q2): https://www.youtube.com/watch?v=M64HUIJFTZM -- And another similar one (though technically Putnam, not IMO): https://www.youtube.com/watch?v=OkmNXy7er84 - Timothy Gowers (Fields Medalist and…

It takes Tim Gowers more than hour and a half to go through q4! (Sure, he could go faster without video. But Tim Gowers! An hour and a half!!)

Re: OpenAI claims gold-medal performance at IMO 2025

#546
post #9

Some previous predictions: In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025. He thought there was an 8% chance of this happening. Eliezer Yudkowsky said "at least 16%". Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...

While I usually enjoy seeing these discussions, I think they are really pushing the usefulness of bayesian statistics. If one dude says the chance for an outcome is 8% and another says it's 16% and the outcome does occur, they were both pretty wrong, even though it might seem like the one who guessed a few % higher might have had a better belief system. Now if one of them had said 90% while the other said 8% or 16%,…

If i predict that my next dice roll will be a 5 with 16% certainty and i do indeed roll a 5, was my prediction wrong?

Re: OpenAI claims gold-medal performance at IMO 2025

#547

Earlier quoted context omitted.

It is not "supremacist" to believe that depriving hundreds of millions of people from higher education in their native language is deeply unjust. This reflection was prompted by a comment on why Indian languages are not represented in international competitions, which was prompted by a comment on the competition being available in many languages. Discussions online have a tendency to go off into tangents like this. I…

Much more efficient for us to all speak the same language. Trying to create fragmentation is inefficient.

You should take that up with the IMO then, or all of European Union. They provide services in ~two dozen languages.

Re: OpenAI claims gold-medal performance at IMO 2025

#548
post #490

Pre-registering a prediction: When (not if) AI does make a major scientific discovery, we'll hear "well it's not really thinking, it just processed all human knowledge and found patterns we missed - that's basically cheating!"

Turns out goalposts are the world’s most easily moved objects. We should start building spacecraft out of them.

Re: OpenAI claims gold-medal performance at IMO 2025

#549
post #526

Earlier quoted context omitted.

Interesting observation. One one hand, these resemble more the notes that an actual participant would write while solving the problem. Also, less words = less noise, more focus. But also, specifically for LLMs that output one token at a time and have a limited token context, I wonder if limiting itself to semantically meaningful tokens can be create longer stretches of semantically coherent thought?

The original thread mentions “test-time compute scaling” so they had some architecture generating a lot of candidate ideas to evaluate. Minimizing tokens can be very meaningful from a scalability perspective alone!

This is just speculation but I wouldn't be surprised if there were some symbolic AI 'tricks'/tools (and/or modern AI trained to imitiate symbolic AI) under the hood.

Re: OpenAI claims gold-medal performance at IMO 2025

#550

Earlier quoted context omitted.

Are you in the US? Have you heard of the AMC (used to be AHMSE) and the AIME? Those are the feeders to the IMO. If your school had a math team and you were on it, would be surprised if you didn't hear of it You may not have heard of the IMO because no one in school district, possibly even state got in. It is extremely selective (like 20 students in the entire country)

I’m in the US but this was a while back, in the south. It was a highly ranked school and ended up producing lots of PhDs, but many of the families were blue collar and so there just wasn’t any awareness of things like this.

I'm curious on your response to GP's question. Have you heard of AHSME, AMC, or AIME?

Nobody mentioned them in high school (1997) until I heard of them online and got my school to participate. 30 kids took the AHSME. Only one qualified for the AIME. And nobody qualified for IMO (though I tell myself I was close).

I believe the 1 in a million number.

Post reply on HN