Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

391–400 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#392
post #9

Some previous predictions: In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025. He thought there was an 8% chance of this happening. Eliezer Yudkowsky said "at least 16%". Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...

Off topic, but am I the only one getting triggered every time I see a rationalist quantify their prediction of the future with single digit accuracy? It's like their magic way of trying to get everyone to forget that they reached their conclusion in completely hand-wavy way, just like every other human being. But instead of saying "low confidence" or "high confidence" like the rest of us normies, they will tell you t…

Would you also get triggered if you saw people make a bet at, say, $24 : $87 odds? Would you shout: "No! That's too precise, you should bet $20 : $90!"? For that matter, should all prices in the stock market be multiples of $1, (since, after all, fluctuations of greater than $1 are very common)?

If the variance (uncertainty) in a number is large, correct thing to do is to just also report the variance, not to round the mean to a whole number.

Also, in log odds, the difference between 5% and 10% is about the same as the difference between 40% and 60%. So using an intermediate value like 8% is less crazy than you'd think.

People writing comments in their own little forum where they happen not to use sig-figs to communicate uncertainty is probably not a sinister attempt to convince "everyone" that their predictions are somehow scientific. For one thing, I doubt most people are dumb enough to be convinced by that, even if it were the goal. For another, the expected audience for these comments was not "everyone", it was specifically people who are likely to interpret those probabilities in a Bayesian way (i.e. as subjective probabilities).

Re: OpenAI claims gold-medal performance at IMO 2025

#393

And of course it's available even in Icelandic, spoken by ~300k people, but not a single Indian language, spoken by hundreds of millions. भारत दुर्दशा न देखी जाई...

Presumably almost all competitors from India would be fluent in English (given it is the second most spoken language there)? I guess the same is true of Icelandic though.

Re: OpenAI claims gold-medal performance at IMO 2025

#394

Has anyone independently reviewed these solutions? My proving skills are extremely rusty so I can’t look at these and validate them. They certainly are not traditional proofs though.

I read through P1, and it seemed to be correct. Though you could explain the central idea of the proof into about 3 sentences and a few drawings. It reads like someone who found the correct answer but seemingly had no understanding of what they did and just handed in the draft paper. Which seems odd, shouldn't an LLM be better at prose?

One would think. I suppose OpenAI threw the majority of their compute budget at producing and verifying solutions. It would certainly be interesting to see whether or not this new model can distill its responses to just those steps necessary to convey its result to a given audience.

Re: OpenAI claims gold-medal performance at IMO 2025

#395
post #308
post #205

Earlier quoted context omitted.

[flagged]

Please don't cross into personal attack. We ban accounts that do that. Also, please don't fulminate. This is in the site guidelines: https://news.ycombinator.com/newsguidelines.html .

Noted.

General attacks are fine, but we draw the line at personal.

Re: OpenAI claims gold-medal performance at IMO 2025

#396
I fed the problem 1 solution into gemini and asked if it was generated by a human or llm. It said:

Conclusion: It is overwhelmingly likely that this document was generated by a human.

----

Self-Correction/Refinement and Explicit Goals:

"Exactly forbidden directions. Good." - This self-affirmation is very human.

"Need contradiction for n>=4." - Clearly stating the goal of a sub-proof.

"So far." - A common human colloquialism in working through a problem.

"Exactly lemma. Good." - Another self-affirmation.

"So main task now: compute K_3. And also show 0,1,3 achievable all n. Then done." - This is a meta-level summary of the remaining work, typical of human problem-solving "

----

Re: OpenAI claims gold-medal performance at IMO 2025

#397

Interesting that the proofs seem to use a limited vocabulary: https://github.com/aw31/openai-imo-2025-proofs/blob/main/pro... Why waste time say lot word when few word do trick :) Also worth pointing out that Alex Wei is himself a gold medalist at IOI.

> Also worth pointing out that Alex Wei is himself a gold medalist at IOI. Terence Tao also called it, that the top LLMs would get gold this year in a recent podcast.

[deleted]

Re: OpenAI claims gold-medal performance at IMO 2025

#398

I fed the problem 1 solution into gemini and asked if it was generated by a human or llm. It said: Conclusion: It is overwhelmingly likely that this document was generated by a human. ---- Self-Correction/Refinement and Explicit Goals: "Exactly forbidden directions. Good." - This self-affirmation is very human. "Need contradiction for n>=4." - Clearly stating the goal of a sub-proof. "So far." - A common human colloq…

LLMs are not an accurate test of whether something was written by an LLM or not.

Re: OpenAI claims gold-medal performance at IMO 2025

#399

Earlier quoted context omitted.

What's the clear path to improved efficiency now that we've reached peak data?

> now that we've reached peak data? A) that's not clear B) now we have "reasoning" models that can be used to analyse the data, create n rollouts for each data piece, and "argue" for / against / neutral on every piece of data going into the model. Imagine having every page of a "short story book" + 10 best "how to write" books, and do n x n on them. Huge compute, but basically infinite data as well. We went from "a b…

A) We are out of the Internet-scale-for-free data. Of course the companies deploying LLM based systems at massive scale are of course ingesting a lot of human data from their users, that they are seeking to use to further improve their models.

B) Has learning though "self-play" (like with AlphaZero etc) been demonstrated working for improving LLMs? What is the latest key research on this?

Re: OpenAI claims gold-medal performance at IMO 2025

#400

Earlier quoted context omitted.

Solving climate change isn't a technical problem, but a human one. We know the steps we have to take, and have for many years. The hard part is getting people to actually do them. No human has any idea how to accomplish that. If a machine could, we would all have much to learn from it.

I disagree with this assessment. We don’t know the steps we have to take. We know a set of steps we could take but they’re societally unpalatable. Technology can potentially offer alternative steps or introduce societal changes that make the first set of steps more palatable.

I feel I should clarify as clearly this is an unpopular opinion: I’m not saying climate change can be solved by technology alone, but I do believe that enabling the societal changes needed to deal with climate change requires using every tool we have at our disposal and that includes technology. I don’t really see why this is controversial and would love to hear that perspective.
Post reply on HN