Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

231–240 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#231
post #79

Earlier quoted context omitted.

I am a professor in a math department (I teach statistics but there is a good complement of actual math PhDs) and there are only about 10% who care about these types of problems and definitely less than half who could get gold on an IMO test even if they didn’t care. They are all outstanding mathematicians, but the IMO type questions are not something that mathematicians can universally solve without preparation. The…

100% agree with this. My second degree is in mathematics. Not only can I probably not do these but they likely aren’t useful to my work so I don’t actually care. I’m not sure an LLM could replace the mathematical side of my work (modelling). Mostly because it’s applied and people don’t know what they are asking for, what is possible or how to do it and all the problems turn out to be quite simple really.

100% agree about this too (also a professional mathematician). To mathematicians who have not been trained on such problems, these will typically look very hard, especially the more recent olympiad problems (as opposed to problems from eg 30 years ago). Basically these problems have become more about mastering a very impressive list of techniques than at the inception (and participants prepare more and more for these). On the other hand, research mathematics has become more and more technical, but the techniques are very different, so that the correlation between olympiads and research is probably smaller than it once was.

Re: OpenAI claims gold-medal performance at IMO 2025

#232

The cynicism/denial on HN about AI is exhausting. Half the comments are some weird form of explaining away the ever increasing performance of these models I've been reading this website for probably 15 years, its never been this bad. many threads are completely unreadable, all the actual educated takes are on X, its almost like there was a talent drain

It’s the same with anything related to cryptocurrency. HN has a hate boner for certain topics.

Re: OpenAI claims gold-medal performance at IMO 2025

#233

My view is that it's less impressive than previous go and chess results. Humans are worse at competitive math than at those games, it's still very limited space and well defined problems. They may hype "general purpose" as much as they want but for now it's still the case that AI is super human at well defined limited space tasks and can't achieve performance of a mediocre below average human at simple tasks without…

The significance, though, is that the "very limited space and well defined problems" continue to expand. Moving from a purpose built system for playing a single game, vs. having a system that can address a broader set of problems would still be a significant step - as more high value tasks will fall into it's competency range. It seems the next big step will be on us to improve eval/feedback systems in less defined problems.

Re: OpenAI claims gold-medal performance at IMO 2025

#234
post #106

I believe this company used to present its results and approach in academic papers with enough details so that it could be reproduced by third parties. Now it is just doing a bunch of tweets?

[flagged]

Different model.

Re: OpenAI claims gold-medal performance at IMO 2025

#235
post #3

Earlier quoted context omitted.

[flagged]

Which would be impressive if we knew those problems weren't in the training data already. I mean it is quite impressive how language models are able to mobilize the knowledge they have been trained on, especially since they are able to retrieve information from sources that may be formatted very differently, with completely different problem statement sentences, different variable names and so on, and really operate…

Even if we accept as a premise that these models are doing "smart retrieval" and not "reasoning" (neither of which are being defined here, nor do I think we can tell from this tweet even if they were), it doesn't really change the impact.

There are many industries for which the vast majority of work done is closer to what I think you mean by "smart retrieval" than what I think you mean by "reasoning." Adult primary care and pediatrics, finance, law, veterinary medicine, software engineering, etc. At least half, if not upwards of 80% of the work in each of these fields is effectively pattern matching to a known set of protocols. They absolutely deal in novel problems as well, but it's not the majority of their work.

Philosophically it might be interesting to ask what "reasoning" means, and how we can assess if the LLMs are doing it. But, practically, the impacts to society will be felt even if all they are doing is retrieval.

Re: OpenAI claims gold-medal performance at IMO 2025

#236

The cynicism/denial on HN about AI is exhausting. Half the comments are some weird form of explaining away the ever increasing performance of these models I've been reading this website for probably 15 years, its never been this bad. many threads are completely unreadable, all the actual educated takes are on X, its almost like there was a talent drain

> I've been reading this website for probably 15 years, its never been this bad... all the actual educated takes are on X Almost every technical comment on HN is wrong (see for example essentially all the discussion of Rust async, in which people keep making up silly claims that Rust maintainers then attempt to patiently explain are wrong). The idea that the "educated" takes are on X though... that's crazy talk.

This is true of every forum and every topic. When you actually know something about the topic you realize 90% of the takes about it are garbage.

But in most other sites the statistic is 99%, so HN is still doing much better than average.

Re: OpenAI claims gold-medal performance at IMO 2025

#237

The cynicism/denial on HN about AI is exhausting. Half the comments are some weird form of explaining away the ever increasing performance of these models I've been reading this website for probably 15 years, its never been this bad. many threads are completely unreadable, all the actual educated takes are on X, its almost like there was a talent drain

> I've been reading this website for probably 15 years, its never been this bad... all the actual educated takes are on X Almost every technical comment on HN is wrong (see for example essentially all the discussion of Rust async, in which people keep making up silly claims that Rust maintainers then attempt to patiently explain are wrong). The idea that the "educated" takes are on X though... that's crazy talk.

No on AI, this is really a fringe environment of relatively uninformed commenters, compared to X. X has its craziness but you can curate your feeds by using lists. Here I can't choose who to follow.

And like said, the researchers themselves are on X, even Gary Marcus is there. ;)

Re: OpenAI claims gold-medal performance at IMO 2025

#238

The cynicism/denial on HN about AI is exhausting. Half the comments are some weird form of explaining away the ever increasing performance of these models I've been reading this website for probably 15 years, its never been this bad. many threads are completely unreadable, all the actual educated takes are on X, its almost like there was a talent drain

HN feels very low signal, since it's populated by people who barely interact with the real world X is higher signal, but very group thinky. It's great if you want to know the trends, but gotta be careful not to jump off the cliff with the lemmings. Highest signal is obviously non digital. Going to meetups, coffee/beers with friends, working with your hands, etc.

it used to be high signal though. you have to wonder if the type of people posting on here is different than it used to be

Re: OpenAI claims gold-medal performance at IMO 2025

#239

Noam Brown: > this isn’t an IMO-specific model. It’s a reasoning LLM that incorporates new experimental general-purpose techniques. > it’s also more efficient [than o1 or o3] with its thinking. And there’s a lot of room to push the test-time compute and efficiency further. > As fast as recent AI progress has been, I fully expect the trend to continue. Importantly, I think we’re close to AI substantially contributing…

I assume there was tool use in the fine tuning?

Re: OpenAI claims gold-medal performance at IMO 2025

#240

The cynicism/denial on HN about AI is exhausting. Half the comments are some weird form of explaining away the ever increasing performance of these models I've been reading this website for probably 15 years, its never been this bad. many threads are completely unreadable, all the actual educated takes are on X, its almost like there was a talent drain

[deleted]
Post reply on HN