Live data from Hacker News

OpenAI claims gold-medal performance at IMO 2025

twitter.com

521–530 of 737 posts

Re: OpenAI claims gold-medal performance at IMO 2025

#521

Earlier quoted context omitted.

I mean, sure people can use this to fool themselves. I think usually the cause of someone fooling themselves is "the will to be fooled" , and not so much that fact that they used precise numbers in the their internal monologue as opposed to verbal buckets like "pretty likely", "very unlikely". But if you estimate 56% it sometimes actually makes a difference, then who am I to argue? Sounds super accurate to me. :) In…

> and not so much that fact that they used precise numbers in the their internal monologue as opposed to verbal buckets like "pretty likely", "very unlikely" I am obviously only talking from my personal anecdotal experience, but having been on a bunch of coffee chat in the last few months with people in the AI safety field in SF, and a lot of them being Lesswrong-ers, I experienced a lot of those discussions with ran…

> It could be just a personal mental weakness with numbers with me that is not general, but looking at my interlocutors emotional reactions to their own numerical predictions I do feel quite strongly that this is a general human trait.

Your feeling is correct; anchoring is a thing, and good LessWrongers (I hope to be in that category) know this and keep track of where their prior and not just posterior probabilities come from: https://en.wikipedia.org/wiki/Anchoring_effect

Probably don't in practice, but should. That "should" is what puts the "less" into "less wrong".

Re: OpenAI claims gold-medal performance at IMO 2025

#522

Earlier quoted context omitted.

> These models cannot reason Not trying to be a smarty pants here, but what do we mean by "reason"? Just to make the point, I'm using Claude to help me code right now. In between prompts, I read HN. It does things for me such as coding up new features, looking at the compile and runtime responses, and then correcting the code. All while I sit here and write with you on HN. It gives me feedback like "lock free message…

I'm having actual coding conversations that I used to only have with senior devs, right now, while browsing HN, and code that does what I asked is being produced. I’m using Opus 4 for coding and there is no way that model demonstrates any reasoning or demonstrates any “intelligence” in my opinion. I’ve been through the having conversations phase etc but doesn’t get you very far, better to read a book. I use these mod…

> It will do something brilliant and another 5 dumb things in the same prompt.

it me

Re: OpenAI claims gold-medal performance at IMO 2025

#523
post #490

Pre-registering a prediction: When (not if) AI does make a major scientific discovery, we'll hear "well it's not really thinking, it just processed all human knowledge and found patterns we missed - that's basically cheating!"

If you want credit for getting predictions right, you have to predict something that has less than 100% probability to happen.

Re: OpenAI claims gold-medal performance at IMO 2025

#524
post #91

Earlier quoted context omitted.

Context? Who are these people and what are these numbers and why shouldn't I assume they're pulled from thin air?

You should basically assume they are pulled from thin air. (Or more precisely, from the brain and world model of the people making the prediction.) The point of giving such estimates is mostly an exercise in getting better at understanding the world, and a way to keep yourself honest by making predictions in advance. If someone else consistently gives higher probabilities to events that ended up happening than you di…

> If someone else consistently gives higher probabilities to events that ended up happening than you did, then that's an indication that there's space for you to improve your prediction ability.

Your inference seems ripe for scams.

For example-- if I find out that a critical mass of participants aren't measuring how many participants are expected to outrank them by random chance, I can organize a simplistic service to charge losers for access to the ostensible "mentors."

I think this happened with the stock market-- you predict how many mutual fund managers would beat the market by random chance for a given period. Then you find that same (small) number of mutual fund managers who beat the market and switched to a more lucrative career of giving speeches about how to beat the market. :)

Re: OpenAI claims gold-medal performance at IMO 2025

#525
post #86

Earlier quoted context omitted.

The key bit here is whether the LLM doing the cherry picking had knowledge of the solution. If it didn't, this is a meaningful result. That's why I'd like more info, but I fear OpenAI is going to try to keep things under wraps.

> If it didn't We kind of have to assume it didn't right? Otherwise bragging about the results makes zero sense and would be outright misleading.

"You really think someone would do that, just go on the internet and tell lies?"

[https://youtube.com/watch?v=YWdD206eSv0]

Re: OpenAI claims gold-medal performance at IMO 2025

#526

Interesting that the proofs seem to use a limited vocabulary: https://github.com/aw31/openai-imo-2025-proofs/blob/main/pro... Why waste time say lot word when few word do trick :) Also worth pointing out that Alex Wei is himself a gold medalist at IOI.

Interesting observation. One one hand, these resemble more the notes that an actual participant would write while solving the problem. Also, less words = less noise, more focus. But also, specifically for LLMs that output one token at a time and have a limited token context, I wonder if limiting itself to semantically meaningful tokens can be create longer stretches of semantically coherent thought?

The original thread mentions “test-time compute scaling” so they had some architecture generating a lot of candidate ideas to evaluate. Minimizing tokens can be very meaningful from a scalability perspective alone!

Re: OpenAI claims gold-medal performance at IMO 2025

#527

This is such an interesting time because the percentage of people who are making predictions about AGI happening on the future are going to drop off and the number of people completely ignoring the term AGI will increase.

That doesn't seem likely because the LLMs haven't really delivered any great products that can cover the money spent and so AGI hype is essentially to keep the money flowing.

Re: OpenAI claims gold-medal performance at IMO 2025

#528

Earlier quoted context omitted.

They funded the entire benchmark and didn’t disclose their involvement. They then proceeded to make use of the benchmark while pretending like they weren’t affiliated with EpochAI. That’s a huge omission and more than enough reason to distrust their claims.

IMO their involvement is only an issue if they gained an advantage on the benchmark by it. If they didn't train on the test set then their gained advantage is minimal and I don't see a big problem with it nor do I see an obligation to disclose. Especially since there is a hold-out set that OpenAI doesn't have access to, which can detect any malfeasance.

It's typically difficult to find direct evidence for bias. That is why rules for conflict of interest and disclosure are strict in research and academia. Crucially, something is a conflict of interest if it could be perceived as a conflict of interest by someone external, so it doesn't matter if you think you could judge fairly, it's important if someone else might doubt you could.

Not disclosing a conflict of interest is generally considered a significant ethics violation, because it reduces trust in the general scientific/research system. Thus OpenAI has become untrustworthy in many people's view irrespective if their involvement with the benchmarks creation affected their results or not.

Re: OpenAI claims gold-medal performance at IMO 2025

#529
post #87

I am neither an optimist nor a pessimist for AI. I would likely be called both by the opposite parties. But the fact that AI / LLM is still rapidly improving is impressive in itself and worth celebrating for. Is it perfect, AGI, ASI? No. Is it useless? Absolutely not. I am just happy the prize is so big for AI that there are enough money involve to push for all the hardware advancement. Foundry, Packaging, Interconne…

But unlike the trillion dollars invested in the broadband internet build out between 1998 and 2008, when this 10 year trillion dollar bubble pops, we won't be left with an enduring and useful piece of infrastructure adding a trillion dollars to the global economy annually.

> ... we won't be left with an enduring and useful piece of infrastructure adding a trillion dollars to the global economy annually.

I'm not drinking the AGI kool-aid but I use LLMs daily. We pay not one but two AI subscriptions at home (including Claude).

It's extremely useful. From translation to proof-reading to synthetizing to expanding on something to writing little dumb functions to helping with spreadsheet formulas to documenting code to writing commit messages to helping find movie names (when I only remember very partially the plot) etc.

How is this not already adding a trillion dollars to the economy?

It's not about the infrastructure: all that counts are the models. They're here to stay. They're not going away.

It's the single biggest time-saver I've ever seen for mundane tasks (and, no, it doesn't write good code: it write shitty pathetic underperforming insecure code... And yet it's still useful for proofs of concept / one-offs / throwaway).

Post reply on HN