Earlier quoted context omitted.
> I think we’re close to AI substantially contributing to scientific discovery. The new "Full Self-Driving next year"?
I know it’s a meme but there actually are fully self driving cars, they make thousands of trips every day in a couple US cities.
OpenAI claims gold-medal performance at IMO 2025
401–410 of 737 posts
Re: OpenAI claims gold-medal performance at IMO 2025
#402Earlier quoted context omitted.
> I think we’re close to AI substantially contributing to scientific discovery. The new "Full Self-Driving next year"?
I know it’s a meme but there actually are fully self driving cars, they make thousands of trips every day in a couple US cities.
Re: OpenAI claims gold-medal performance at IMO 2025
#403From that thread: "The model solved P1 through P5; it did not produce a solution for P6." It's interesting that it didn't solve the problem that was by far the hardest for humans too. China, the #1 team got only 21/42 points on it. In most other teams nobody solved it.
To me, this is a tell of human-involvement in the model solution. There is no reason why machines would do badly on exactly the problem which humans do badly as well - without humans prodding the machine towards a solution. Also, there is no reason why machines could not produce a partial or wrong answer to problem 6 which seems like survivor bias to me. ie, that only correct solutions were cherrypicked.
There are many things that are hard for AI’s for the same reason they’re hard for humans. There are subtleties in complexity that make challenging things universal.
Obviously the model was trained on human data so its competencies lie in what other humans have provided input for over the years in mathematics, but that isn’t data contamination, that’s how all humans learn. This model, like the contestants, never saw the questions before.
Re: OpenAI claims gold-medal performance at IMO 2025
#404Re: OpenAI claims gold-medal performance at IMO 2025
#405Wow. That's an impressive result, but how did they do it? Wei references scaling up test-time compute, so I have to assume they threw a boatload of money at this. I've heard talk of running models in parallel and comparing results - if OpenAI ran this 10000 times in parallel and cherry-picked the best one, this is a lot less exciting. If this is legit, then we need to know what tools were used and how the model used…
This is almost certainly the case, remember the initial o3 ARC benchmark? I could add this is probably multi-agent system as well, so the context length restriction can be bypassed.
Overall, AI good at math problems doesn't make news to me. It is already better than 99.99% of humans, now it is better than 99.999% of us. So ... ?
Re: OpenAI claims gold-medal performance at IMO 2025
#406Some previous predictions: In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025. He thought there was an 8% chance of this happening. Eliezer Yudkowsky said "at least 16%". Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...
Re: OpenAI claims gold-medal performance at IMO 2025
#407Some previous predictions: In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025. He thought there was an 8% chance of this happening. Eliezer Yudkowsky said "at least 16%". Source: https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...
Context? Who are these people and what are these numbers and why shouldn't I assume they're pulled from thin air?
The point of giving such estimates is mostly an exercise in getting better at understanding the world, and a way to keep yourself honest by making predictions in advance. If someone else consistently gives higher probabilities to events that ended up happening than you did, then that's an indication that there's space for you to improve your prediction ability. (The quantitative way to compare these things is to see who has lower log loss [1].)
Re: OpenAI claims gold-medal performance at IMO 2025
#408From that thread: "The model solved P1 through P5; it did not produce a solution for P6." It's interesting that it didn't solve the problem that was by far the hardest for humans too. China, the #1 team got only 21/42 points on it. In most other teams nobody solved it.
To me, this is a tell of human-involvement in the model solution. There is no reason why machines would do badly on exactly the problem which humans do badly as well - without humans prodding the machine towards a solution. Also, there is no reason why machines could not produce a partial or wrong answer to problem 6 which seems like survivor bias to me. ie, that only correct solutions were cherrypicked.
Re: OpenAI claims gold-medal performance at IMO 2025
#409Earlier quoted context omitted.
Off topic, but am I the only one getting triggered every time I see a rationalist quantify their prediction of the future with single digit accuracy? It's like their magic way of trying to get everyone to forget that they reached their conclusion in completely hand-wavy way, just like every other human being. But instead of saying "low confidence" or "high confidence" like the rest of us normies, they will tell you t…
Would you also get triggered if you saw people make a bet at, say, $24 : $87 odds? Would you shout: "No! That's too precise, you should bet $20 : $90!"? For that matter, should all prices in the stock market be multiples of $1, (since, after all, fluctuations of greater than $1 are very common)? If the variance (uncertainty) in a number is large, correct thing to do is to just also report the variance, not to round t…
I really wonder what you mean by this. If I put my finger in the air and estimate the emergence of AGI as 13%, how do I get at the variance of that estimate? At face value, it is a number, not a random variable, and does not have a variance. If you instead view it as a "random sample" from the population of possible estimates I might have made, it does not seem well defined at all.
Re: OpenAI claims gold-medal performance at IMO 2025
#410Earlier quoted context omitted.
I know it’s a meme but there actually are fully self driving cars, they make thousands of trips every day in a couple US cities.
> in a couple US cities FWIW, when you get this reductive with your criterion there were technically self-driving cars in 2008 too.