Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

471–480 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#471
post #345

Noam Brown: > this isn’t an IMO-specific model. It’s a reasoning LLM that incorporates new experimental general-purpose techniques. > it’s also more efficient [than o1 or o3] with its thinking. And there’s a lot of room to push the test-time compute and efficiency further. > As fast as recent AI progress has been, I fully expect the trend to continue. Importantly, I think we’re close to AI substantially contributing…

> I think we’re close to AI substantially contributing to scientific discovery. The new "Full Self-Driving next year"?

As an aside, that is happening in China right now in commercial vehicles. I rode a robotaxi last month in Beijing, and those services are expanding throughout China. Really impressive.

Re: OpenAI claims gold-medal performance at IMO 2025

#473
post #215

Earlier quoted context omitted.

With regard to AI & LLMs Twitter/x is actually the only place with all of the industry people discussing. There are a bunch of great accounts to follow that are only really posting content to x. Karpathy, nearcyan, kalomaze, all of the OpenAI researchers including the link this discussion is on, many anthropic researchers. It's such a meme that you see people discuss reading Twitter thread + paper because the thread…

I'm unconvinced Twitter is a very good medium for serious technical discussion. I imagine a lot of this happens on the sidelines at conferences, on mailing threads and actually in organisations doing work on AI (e.g. Universities, Anthropic). The people who are doing the work are also often not the people who have time to Twitter.

Have you published in ML conferences? I'm curious because I have ML researcher friends who have and they talk about Twitter a lot but I'm not an ML researcher myself.

Re: OpenAI claims gold-medal performance at IMO 2025

#474
post #85

I encourage anyone who thinks these are easy high-school problems to try to solve some. They're published (including this year's) at https://www.imo-official.org/problems.aspx . They make my head spin.

Related — these videos give a sense of how someone might actually go about thinking through and solving these kinds of problems:

- A 3Blue1Brown video on a particularly nice and unexpectedly difficult IMO problem (2011 IMO, Q2): https://www.youtube.com/watch?v=M64HUIJFTZM

-- And another similar one (though technically Putnam, not IMO): https://www.youtube.com/watch?v=OkmNXy7er84

- Timothy Gowers (Fields Medalist and IMO perfect scorer) solving this year’s IMO problems in “real time”:

-- Q1: https://www.youtube.com/watch?v=1G1nySyVs2w

-- Q4: https://www.youtube.com/watch?v=O-vp4zGzwIs

Re: OpenAI claims gold-medal performance at IMO 2025

#475
post #335

Earlier quoted context omitted.

> dismissing any claim from employees of large tech companies Me: I have a way to turn lead into gold. You: Show me!!! Me: NO (and then spends the rest of my life in poverty). Cold Fusion (physics not the programing language) is the best example of why you "Show your work". This is the Valley we're talking about. It's the thudnderdome of technology and companies. If you have a meaningful breakthrough you don't talk a…

I don't think this is a reasonable take. Some people/organizations send signals about things that we're not ready to fully drop it on the world. Others consider those signals in context (reputation of sender, prior probability of being true, reasons for sender to be honest vs. deceptive, etc). When my wife tells me there's a pie in the oven and it's smelling particularly good, I don't demand evidence or disbelieve th…

> Some people/organizations send signals about things that we're not ready to fully drop it on the world.

This is called marketing.

> When my wife tells me there's a pie in the oven and it's smelling particularly good, I don't demand evidence

Because you have evidence, it smells.

And if later your ask your wife "where is the pie" and she says "I sprayed pie scent in the air, I was just singling" how are you going to feel?

Open AI spent its "fool us once" card already. Doing things this way does not earn back trust, failure to deliver (and they have done that more than once) ... See staff non disparagement, see the math fiasco, see open weights.

Re: OpenAI claims gold-medal performance at IMO 2025

#476

Earlier quoted context omitted.

I think the main hesitancy is due to rampant anthropomorphism. These models cannot reason, they pattern match language tokens and generate emergent behaviour as a result. Certainly the emergent behaviour is exciting but we tend to jump to conclusions as to what it implies. This means we are far more trusting with software that lacks formal guarantees than we should be. We are used to software being sound by default b…

> These models cannot reason Not trying to be a smarty pants here, but what do we mean by "reason"? Just to make the point, I'm using Claude to help me code right now. In between prompts, I read HN. It does things for me such as coding up new features, looking at the compile and runtime responses, and then correcting the code. All while I sit here and write with you on HN. It gives me feedback like "lock free message…

I think the biggest hint that the models aren't reasoning is that they can't explain their reasoning. Researchers have shown for explained that how a model solves a simple math problem and how it claims to have solved it after the fact have no real correlation. In other words there was only the appearance of reasoning.

Re: OpenAI claims gold-medal performance at IMO 2025

#477
post #433

Earlier quoted context omitted.

Yes, and there's also languages of ex-USSR countries, whose competitors presumably all understand Russian, and so on. The real reason might be that there's an enormous class of self-loathing elites in India who actively despise the possibility of any Indian language being represented in higher education. This obviously stunts the possibility of them being used in international competitions.

Discussions about Indian politics or the Indian psyche—especially when laced with Indic supremacist undertones—are off-topic and an annoyance here. Please consider sharing these views in a forum focused on Indian affairs, where they’re more likely to find the traction they deserve.

It is not "supremacist" to believe that depriving hundreds of millions of people from higher education in their native language is deeply unjust. This reflection was prompted by a comment on why Indian languages are not represented in international competitions, which was prompted by a comment on the competition being available in many languages.

Discussions online have a tendency to go off into tangents like this. It's regrettable that this is such a contentious topic.

Re: OpenAI claims gold-medal performance at IMO 2025

#478

Earlier quoted context omitted.

There is no honor in hiding behind euphemisms. Rationalists say ‘low confidence’ and ‘high confidence’ all the time, just not when they're making an actual bet and need to directly compare credences. And the 16.27% mockery is completely dishonest. They used less than a single significant figure.

> just not when they're making an actual bet That is not my experience talking with rationalists irl at all. And that is precisely my issue, it is pervasive in every day discussion about any topic, at least with the subset of rationalists I happen to cross paths with. If it was just for comparing ability to forecast or for bets, then sure it would make total sense. Just the other day I had a conversation with someone…

I wonder if what you observe is a direct effect of the rationalist movement worshipping the god of Bayes.

Re: OpenAI claims gold-medal performance at IMO 2025

#479

Earlier quoted context omitted.

Off topic, but am I the only one getting triggered every time I see a rationalist quantify their prediction of the future with single digit accuracy? It's like their magic way of trying to get everyone to forget that they reached their conclusion in completely hand-wavy way, just like every other human being. But instead of saying "low confidence" or "high confidence" like the rest of us normies, they will tell you t…

Interestingly, this is actually a question that's been looked at empirically! Take a look at this paper: https://scholar.harvard.edu/files/rzeckhauser/files/value_of... They took high-precision forecasts from a forecasting tournament and rounded them to coarser buckets (nearest 5%, nearest 10%, nearest 33%), to see if the precision was actually conveying any real information. What they found is that if you rounded th…

Aim small, miss small?
Post reply on HN