OpenAI claims gold-medal performance at IMO 2025
391–400 of 737 posts
Re: OpenAI claims gold-medal performance at IMO 2025
#392Some previous predictions: In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025. He thought there was an 8% chance of this happening. Eliezer Yudkowsky said "at least 16%". Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...
Off topic, but am I the only one getting triggered every time I see a rationalist quantify their prediction of the future with single digit accuracy? It's like their magic way of trying to get everyone to forget that they reached their conclusion in completely hand-wavy way, just like every other human being. But instead of saying "low confidence" or "high confidence" like the rest of us normies, they will tell you t…
If the variance (uncertainty) in a number is large, correct thing to do is to just also report the variance, not to round the mean to a whole number.
Also, in log odds, the difference between 5% and 10% is about the same as the difference between 40% and 60%. So using an intermediate value like 8% is less crazy than you'd think.
People writing comments in their own little forum where they happen not to use sig-figs to communicate uncertainty is probably not a sinister attempt to convince "everyone" that their predictions are somehow scientific. For one thing, I doubt most people are dumb enough to be convinced by that, even if it were the goal. For another, the expected audience for these comments was not "everyone", it was specifically people who are likely to interpret those probabilities in a Bayesian way (i.e. as subjective probabilities).
Re: OpenAI claims gold-medal performance at IMO 2025
#393And of course it's available even in Icelandic, spoken by ~300k people, but not a single Indian language, spoken by hundreds of millions. भारत दुर्दशा न देखी जाई...
Re: OpenAI claims gold-medal performance at IMO 2025
#394Has anyone independently reviewed these solutions? My proving skills are extremely rusty so I can’t look at these and validate them. They certainly are not traditional proofs though.
I read through P1, and it seemed to be correct. Though you could explain the central idea of the proof into about 3 sentences and a few drawings. It reads like someone who found the correct answer but seemingly had no understanding of what they did and just handed in the draft paper. Which seems odd, shouldn't an LLM be better at prose?
Re: OpenAI claims gold-medal performance at IMO 2025
#395Earlier quoted context omitted.
[flagged]
Please don't cross into personal attack. We ban accounts that do that. Also, please don't fulminate. This is in the site guidelines: https://news.ycombinator.com/newsguidelines.html .
General attacks are fine, but we draw the line at personal.
Re: OpenAI claims gold-medal performance at IMO 2025
#396Conclusion: It is overwhelmingly likely that this document was generated by a human.
----
Self-Correction/Refinement and Explicit Goals:
"Exactly forbidden directions. Good." - This self-affirmation is very human.
"Need contradiction for n>=4." - Clearly stating the goal of a sub-proof.
"So far." - A common human colloquialism in working through a problem.
"Exactly lemma. Good." - Another self-affirmation.
"So main task now: compute K_3. And also show 0,1,3 achievable all n. Then done." - This is a meta-level summary of the remaining work, typical of human problem-solving "
----
Re: OpenAI claims gold-medal performance at IMO 2025
#397Interesting that the proofs seem to use a limited vocabulary: https://github.com/aw31/openai-imo-2025-proofs/blob/main/pro... Why waste time say lot word when few word do trick :) Also worth pointing out that Alex Wei is himself a gold medalist at IOI.
> Also worth pointing out that Alex Wei is himself a gold medalist at IOI. Terence Tao also called it, that the top LLMs would get gold this year in a recent podcast.
Re: OpenAI claims gold-medal performance at IMO 2025
#398I fed the problem 1 solution into gemini and asked if it was generated by a human or llm. It said: Conclusion: It is overwhelmingly likely that this document was generated by a human. ---- Self-Correction/Refinement and Explicit Goals: "Exactly forbidden directions. Good." - This self-affirmation is very human. "Need contradiction for n>=4." - Clearly stating the goal of a sub-proof. "So far." - A common human colloq…
Re: OpenAI claims gold-medal performance at IMO 2025
#399Earlier quoted context omitted.
What's the clear path to improved efficiency now that we've reached peak data?
> now that we've reached peak data? A) that's not clear B) now we have "reasoning" models that can be used to analyse the data, create n rollouts for each data piece, and "argue" for / against / neutral on every piece of data going into the model. Imagine having every page of a "short story book" + 10 best "how to write" books, and do n x n on them. Huge compute, but basically infinite data as well. We went from "a b…
B) Has learning though "self-play" (like with AlphaZero etc) been demonstrated working for improving LLMs? What is the latest key research on this?
Re: OpenAI claims gold-medal performance at IMO 2025
#400Earlier quoted context omitted.
Solving climate change isn't a technical problem, but a human one. We know the steps we have to take, and have for many years. The hard part is getting people to actually do them. No human has any idea how to accomplish that. If a machine could, we would all have much to learn from it.
I disagree with this assessment. We don’t know the steps we have to take. We know a set of steps we could take but they’re societally unpalatable. Technology can potentially offer alternative steps or introduce societal changes that make the first set of steps more palatable.