Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

701–710 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#701
post #555

Earlier quoted context omitted.

In context, the "at least 16%" is responding to someone who said 8%, and 16 just happens to be exactly twice 8. I suspect (though I don't know) that Yudkowsky would not have claimed to have a robust way to pick whether 16% or 17% was the better figure. For what it's worth, I don't think there's anything even slightly wrong with using whatever estimate feels good to you, even if it happens not to fit someone else's cr…

> In context, the "at least 16%" is responding to someone who said 8%, and 16 just happens to be exactly twice 8. I suspect (though I don't know) that Yudkowsky would not have claimed to have a robust way to pick whether 16% or 17% was the better figure. If this was just a way to say "at least double that", that's... fair enough, I guess. Regarding your other point: > For what it's worth, I don't think there's anythi…

I'm not missing the point, I'm disagareeing with it. I am saying that the convention that if you say 20% then you are assumed to have an error margin of 5%, while if you say 19% you are assumed to have an error margin of 1%, is a bad convention. It gives you no way to say that the number is 20% with a margin of 1%. It gives you only a very small set of possible degrees-of-uncertainty. It gives you no way to express that actually your best estimate is somewhat below 20% even though you aren't sure it isn't 5% out.

It's true, of course, that if you are talking to people who are going to interpret "20%" as "anywhere between 17.5% and 22.5%" and "19%" as "anywhere between 18.5% and 19.5%", then you should try to avoid giving not-round numbers when your uncertainty is high. And that many people do interpret things that way, because although I think the convention is a bad one it's certainly a common one.

But: that isn't what happened in the case you're complaining about. It was a discussion on Less Wrong, where all the internet-rationalists hang out, and where there is not a convention that giving a not-round number implies high confidence and high precision. Also, I looked up what Yudkowsky actually wrote, and it makes it perfectly clear (explicitly, rather than via convention) that his level of uncertainty was high:

"Ha! Okay then. My probability is at least 16%, though I'd have to think more and Look into Things, and maybe ask for such sad little metrics as are available before I was confident saying how much more."

(Incidentally, in case anyone's similarly salty about the 8% figure that gives context to this one: it wasn't any individual's estimate, it was a Metaculus prediction, and it seems pretty obvious to me that it is not an improvement to report a Metaculus prediction of 8% as "a little under 10%" or whatever.)

Re: OpenAI claims gold-medal performance at IMO 2025

#703

Earlier quoted context omitted.

>Why is that less exciting? Because if I have to throw 10000 rocks to get one in the bucket, I am not as good/useful of a rock-into-bucket-thrower as someone who gets it in one shot. People would probably not be as excited about the prospect of employing me to throw rocks for them.

It’s exciting because nearly all humans have 0% chance of throwing the rock into the bucket, and most people believed a rock-into-bucket-thrower machine is impossible. So even an inefficient rock-into-bucket-thrower is impressive. But the bar has been getting raised very rapidly. What was impressive six months ago is awful and unexciting today.

You're putting words in my mouth. It's not "awful and unexciting", it is certainly an important step, but the hype being invited with the headline is the immensely greater one of an accurate rock-thrower. And if they have the inefficient one and trying to pretend to have the real deal, that's flim-flam-man levels of overstatement.

Re: OpenAI claims gold-medal performance at IMO 2025

#704

Earlier quoted context omitted.

I think the biggest hint that the models aren't reasoning is that they can't explain their reasoning. Researchers have shown for explained that how a model solves a simple math problem and how it claims to have solved it after the fact have no real correlation. In other words there was only the appearance of reasoning.

Is this true though? I've suggested things that it pushed back on. Feels very much like a dev. It doesn't just dumbly do what I tell it.

Sure but it isn't reasoning that it should push back. It isn't even "pushing" which would require an intent to change you which it lacks

Re: OpenAI claims gold-medal performance at IMO 2025

#705

Earlier quoted context omitted.

I think the biggest hint that the models aren't reasoning is that they can't explain their reasoning. Researchers have shown for explained that how a model solves a simple math problem and how it claims to have solved it after the fact have no real correlation. In other words there was only the appearance of reasoning.

People can't explain their reasoning either. People do a parallel construction of logical arguments for a conclusion they already reached intuitively in a way they have no clue how it happened. "The idea just popped into my head while showering" to our credit, if this post-hoc rationalization fails we are able to change our opinion to some degree.

Yeah, surprisingly I think the differences are less in the mechanism used for thought and more in the experience of being a person alive in a body. A person can become an idea. An LLM always forgets everything. It cannot "care"

Re: OpenAI claims gold-medal performance at IMO 2025

#706

Earlier quoted context omitted.

People can't explain their reasoning either. People do a parallel construction of logical arguments for a conclusion they already reached intuitively in a way they have no clue how it happened. "The idea just popped into my head while showering" to our credit, if this post-hoc rationalization fails we are able to change our opinion to some degree.

Yeah, surprisingly I think the differences are less in the mechanism used for thought and more in the experience of being a person alive in a body. A person can become an idea. An LLM always forgets everything. It cannot "care"

[dead]

Re: OpenAI claims gold-medal performance at IMO 2025

#707
post #694

Earlier quoted context omitted.

That's just the frequentist approach. But we're talking about bayesian statistics here.

I admit I dont know Bayesian, but isn't the only way to check if the future teller is lucky or not to have them predict many things? If he predicts 10 to happen with a 10% chance, and one of them happens, he's good. If he predicts 10 to happen with a 90% chance and 9 happen, same. How is this different with Bayesian?

It is the only way if you're a frequentist. But there is a whole other subfield of statistics that deals with assigning probabilities to single events.

Re: OpenAI claims gold-medal performance at IMO 2025

#708

Earlier quoted context omitted.

While I usually enjoy seeing these discussions, I think they are really pushing the usefulness of bayesian statistics. If one dude says the chance for an outcome is 8% and another says it's 16% and the outcome does occur, they were both pretty wrong, even though it might seem like the one who guessed a few % higher might have had a better belief system. Now if one of them had said 90% while the other said 8% or 16%,…

The person who guessed 16% would have a lower Brier score (lower is better) and someone who estimated 100%, beyond being correct, would have the lowest possible value.

I'm not saying there aren't ways to measure this (bayesian statistics does exist after all), I'm saying the difference is not worth arguing about who was right. Or even who had a better guess.

Re: OpenAI claims gold-medal performance at IMO 2025

#709
post #85

I encourage anyone who thinks these are easy high-school problems to try to solve some. They're published (including this year's) at https://www.imo-official.org/problems.aspx . They make my head spin.

Related — these videos give a sense of how someone might actually go about thinking through and solving these kinds of problems: - A 3Blue1Brown video on a particularly nice and unexpectedly difficult IMO problem (2011 IMO, Q2): https://www.youtube.com/watch?v=M64HUIJFTZM -- And another similar one (though technically Putnam, not IMO): https://www.youtube.com/watch?v=OkmNXy7er84 - Timothy Gowers (Fields Medalist and…

For people who prefer reading to watching videos, I wrote a detailed account of my process for solving one of last year's IMO problems, along with thoughts on how this relates to AI:

https://secondthoughts.ai/p/solving-math-olympiad-problems

Re: OpenAI claims gold-medal performance at IMO 2025

#710

Earlier quoted context omitted.

I'd disagree with this take. Math olympiads are some of the most intellectually creative activities I've ever done that fit within a one day time limit. Chess and go don't even come close--I am not a strong player, but I've studied both games for hundreds of hours. (My hot take is that chess is not even very creative at all, that's why classical AI techniques produced super human results many years ago.) There is no…

What you are missing about chess and go is that those games are not about finding one true solution. They are very psychological games (at human level) and are about finding moves that are difficult to handle for the opponent. You try to understand how your opponent thinks and what is going to be unpleasant for them. This gives a lot of scope for creative and psychological warfare. In competitive math (or programming…

> They are very psychological games (at human level) and are about finding moves that are difficult to handle for the opponent. You try to understand how your opponent thinks and what is going to be unpleasant for them. This gives a lot of scope for creative and psychological warfare.

And yet even basic models which can run on my phone win this psychological warfare with best players in the world. The scope of problems on IMO is unlimited. Please note that IMO is won by literally best high-school students in the world and most of them are unable to solve all problems (even gold medal winners). Do you think that they are dumb and unable to learn "few tricks"?

>In competitive math (or programming) there is one correct solution and no opponent. It's just not possible for it to be very creative endeavor if those solutions can be found in very limited time.

That's absurd. You could say same things about math research (and "one correct solution" would be wrong as it is for IMO), do you consider it something that's not creative?

Post reply on HN