Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

431–440 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#431
post #7

Progress is astounding. Recently report published about evaluation of LLMs on IMO 2025. o3 high didn't even get bronze. https://matharena.ai/imo/ Waiting for Terry Tao's thoughts, but these kind of things are good use of AI. We need to make science progress faster rather than disrupting our economy without being ready.

[flagged]

Please see https://news.ycombinator.com/item?id=44617609.

You degraded this thread badly by posting so many comments like this.

Re: OpenAI claims gold-medal performance at IMO 2025

#432

Earlier quoted context omitted.

Off topic, but am I the only one getting triggered every time I see a rationalist quantify their prediction of the future with single digit accuracy? It's like their magic way of trying to get everyone to forget that they reached their conclusion in completely hand-wavy way, just like every other human being. But instead of saying "low confidence" or "high confidence" like the rest of us normies, they will tell you t…

No, you are right, this hyper-numericalism is just astrology for nerds.

The whole community is very questionable, at best. (AI 2027, etc.)

Re: OpenAI claims gold-medal performance at IMO 2025

#433

Earlier quoted context omitted.

Presumably almost all competitors from India would be fluent in English (given it is the second most spoken language there)? I guess the same is true of Icelandic though.

Yes, and there's also languages of ex-USSR countries, whose competitors presumably all understand Russian, and so on. The real reason might be that there's an enormous class of self-loathing elites in India who actively despise the possibility of any Indian language being represented in higher education. This obviously stunts the possibility of them being used in international competitions.

Discussions about Indian politics or the Indian psyche—especially when laced with Indic supremacist undertones—are off-topic and an annoyance here. Please consider sharing these views in a forum focused on Indian affairs, where they’re more likely to find the traction they deserve.

Re: OpenAI claims gold-medal performance at IMO 2025

#434
post #425

And of course it's available even in Icelandic, spoken by ~300k people, but not a single Indian language, spoken by hundreds of millions. भारत दुर्दशा न देखी जाई...

Please don't take HN threads into nationalistic flamewar. It leads nowhere interesting or good. We detached this subthread from https://news.ycombinator.com/item?id=44615783 .

[flagged]

Re: OpenAI claims gold-medal performance at IMO 2025

#435
post #273

[dead]

> such announcements should wait at least a week after the closing ceremony it would raise more concerns that corps leaked questions/answers to training data and finetuned specialized models during this time.

This is solvable? They could post a hash of the solutions publicly, and reveal the contents of the solution after one week.

Re: OpenAI claims gold-medal performance at IMO 2025

#436

Earlier quoted context omitted.

And yet when working on production code current LLMs are about as good as a poor intern. Not sure why the disconnect.

Depends. I’ve been using it for some of my workflows and I’d say it is more like a solid junior developer with weird quirks where it makes stupid mistakes and other times behaves as a 30 year SME vet.

I really doubt it's like a "solid junior developer". If it could do the work of a solid junior developer it would be making programming projects 10-100x faster because it can do things several times faster than a person can. Maybe it can write solid code for certain tasks but that's not the same thing as being a junior developer.

Re: OpenAI claims gold-medal performance at IMO 2025

#437

Earlier quoted context omitted.

> now that we've reached peak data? A) that's not clear B) now we have "reasoning" models that can be used to analyse the data, create n rollouts for each data piece, and "argue" for / against / neutral on every piece of data going into the model. Imagine having every page of a "short story book" + 10 best "how to write" books, and do n x n on them. Huge compute, but basically infinite data as well. We went from "a b…

A) We are out of the Internet-scale-for-free data. Of course the companies deploying LLM based systems at massive scale are of course ingesting a lot of human data from their users, that they are seeking to use to further improve their models. B) Has learning though "self-play" (like with AlphaZero etc) been demonstrated working for improving LLMs? What is the latest key research on this?

Certainly the models have orders of magnitude more data available to them than the smartest human being who ever lived does/did. So we can assume that if the goal is "merely" superhuman intelligence, data is not a problem.

It might be a constraint on the evolution of godlike intelligence, or AGI. But at that point we're so far out in bong-hit territory that it will be impossible to say who's right or wrong about what's coming.

Has learning though "self-play" (like with AlphaZero etc) been demonstrated working for improving LLMs?

My understanding (which might be incorrect) is that this amounts to RLHF without the HF part, and is basically how DeepSeek-R1 was trained. I recall reading about OpenAI being butthurt^H^H^H^H^H^H^H^H concerned that their API might have been abused by the Chinese to train their own model.

Re: OpenAI claims gold-medal performance at IMO 2025

#438

Earlier quoted context omitted.

That's a big leap from "answering test questions" to "contributing to scientific discovery".

Having spent tens of thousands of hours contributing to scientific discovery by reading dense papers for a single piece of information, reverse engineering code written by biologists, and tweaking graphics to meet journal requirements… I can say with certainty it’s already contributing by allowing scientists to spend time on science versus spending an afternoon figuring out which undocumented argument in a R package…

This. Even if LLM’s ultimately hit some hard ceiling as substantially-better-Googling-automatons they would already accelerate all thought-based work across the board, and that’s the level they’re already at now (arguably they’re beyond that).

We’re already at the point where these tools are removing repetitive/predictable tasks from researchers (and everyone else), so clearly they’re already accelerating research.

Re: OpenAI claims gold-medal performance at IMO 2025

#439
post #160

Earlier quoted context omitted.

A third party tried this experiment with publicly available models. OpenAI did half as well as Gemini, and none of the models even got bronze. https://matharena.ai/imo/

I feel you're misunderstanding something. That's not "this exact experiment". Matharena is testing publicly available models against the IMO problem set. OpenAI was announcing the results of a new, unpublished model, on that problems set. It is totally fair to discount OpenAI's statement until we have way more details about their setup, and maybe even until there is some level of public access to the model. But you'r…

If OpenAI would publish the models before the competition, then one could verify that they were not tinkered with. Assuming that there exists a way for them to prove that a model is the same, at least. Since the weights are not open, the most basic approach is void.

Re: OpenAI claims gold-medal performance at IMO 2025

#440

Earlier quoted context omitted.

Off topic, but am I the only one getting triggered every time I see a rationalist quantify their prediction of the future with single digit accuracy? It's like their magic way of trying to get everyone to forget that they reached their conclusion in completely hand-wavy way, just like every other human being. But instead of saying "low confidence" or "high confidence" like the rest of us normies, they will tell you t…

Interestingly, this is actually a question that's been looked at empirically! Take a look at this paper: https://scholar.harvard.edu/files/rzeckhauser/files/value_of... They took high-precision forecasts from a forecasting tournament and rounded them to coarser buckets (nearest 5%, nearest 10%, nearest 33%), to see if the precision was actually conveying any real information. What they found is that if you rounded th…

Likely vs. unlikely is rounding to 50%. Single digit is rounding to 1%. I don't think the parent was suggesting the former is better than the latter. Even before I read your comment I thought that 5% precision is useful but 1% precision is a silly turn-off, unless that 1% is near the 0% or 100% boundary.
Post reply on HN