This is also wildly ahead in SWE-bench (71.7%, previous 48%) and Frontier Math (25% on high compute, previous 2%). So much for a plateau lol.
OpenAI O3 breakthrough high score on ARC-AGI-PUB
11–20 of 1001 posts
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#12If people constantly have to ask if your test is a measure of AGI, maybe it should be renamed to something else.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#13A lot of people have criticized ARC as not being relevant or indicative of true reasoning, but I think it was exactly the right thing. The fact that scaled reasoning models are finally showing progress on ARC proves that what it measures really is relevant and important for reasoning.
It's obvious to everyone that these models can't perform as well as humans on everyday tasks despite blowout scores on the hardest tests we give to humans. Yet nobody could quantify exactly the ways the models were deficient. ARC is the best effort in that direction so far.
We don't need more "hard" benchmarks. What we need right now are "easy" benchmarks that these models nevertheless fail. I hope Francois has something good cooked up for ARC 2!
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#14Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#15Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#16How much longer can I get paid $150k to write code ?
This I think will happen with programmers. Rote programming will slowly die out, while demand for super high end will go dramatically up in price.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#17"low compute" mode: Uses 6 samples per task, Uses 33M tokens for the semi-private eval set, Costs $17-20 per task, Achieves 75.7% accuracy on semi-private eval
The "high compute" mode: Uses 1024 samples per task (172x more compute), Cost data was withheld at OpenAI's request, Achieves 87.5% accuracy on semi-private eval
Can we just extrapolate $3kish per task on high compute? (wondering if they're withheld because this isn't the case?)
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#18This is also wildly ahead in SWE-bench (71.7%, previous 48%) and Frontier Math (25% on high compute, previous 2%). So much for a plateau lol.
It’s been really interesting to watch all the internet pundits’ takes on the plateau… as if the two years since the release of GPT3.5 is somehow enough data for an armchair ponce to predict the performance characteristics of an entirely novel technology that no one understands.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#19My skeptical impression: it's complete hubris to conflate ARC or any benchmark with truly general intelligence.
I know my skepticism here is identical to moving goalposts. More and more I am shifting my personal understanding of general intelligence as a phenomenon we will only ever be able to identify with the benefit of substantial retrospect.
As it is with any sufficiently complex program, if you could discern the result beforehand, you wouldn't have had to execute the program in the first place.
I'm not trying to be a downer on the 12th day of Christmas. Perhaps because my first instinct is childlike excitement, I'm trying to temper it with a little reason.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#20So now not only are the models closed, but so are their evals?! This is a "semi-private" eval. WTH is that supposed to mean? I'm sure the model is great but I refuse to take their word for it.