With only a 100x increase in cost, we improved performance by 0.1x and continued plotting this concave-down diminishing-returns type graph! Hurray for logarithmic x-axes! Joking aside, better than ever before at any cost is an achievement, it just doesn't exactly scream "breakthrough" to me.
OpenAI O3 breakthrough high score on ARC-AGI-PUB
521–530 of 1001 posts
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#522With only a 100x increase in cost, we improved performance by 0.1x and continued plotting this concave-down diminishing-returns type graph! Hurray for logarithmic x-axes! Joking aside, better than ever before at any cost is an achievement, it just doesn't exactly scream "breakthrough" to me.
It may eventually be able to solve any problem
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#523With only a 100x increase in cost, we improved performance by 0.1x and continued plotting this concave-down diminishing-returns type graph! Hurray for logarithmic x-axes! Joking aside, better than ever before at any cost is an achievement, it just doesn't exactly scream "breakthrough" to me.
[flagged]
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#524With only a 100x increase in cost, we improved performance by 0.1x and continued plotting this concave-down diminishing-returns type graph! Hurray for logarithmic x-axes! Joking aside, better than ever before at any cost is an achievement, it just doesn't exactly scream "breakthrough" to me.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#525Direct quote from the ARC-AGI blog: “SO IS IT AGI? ARC-AGI serves as a critical benchmark for detecting such breakthroughs, highlighting generalization power in a way that saturated or less demanding benchmarks cannot. However, it is important to note that ARC-AGI is not an acid test for AGI – as we've repeated dozens of times this year. It's a research tool designed to focus attention on the most challenging unsolve…
> acid test The css acid test? This can be gamed too.
> An acid test is a qualitative chemical or metallurgical assay utilizing acid. Historically, it often involved the use of a robust acid to distinguish gold from base metals. Figuratively, the term represents any definitive test for attributes, such as gauging a person's character or evaluating a product's performance.
Specifically here, they're using the figurative sense of "definitive test".
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#526Earlier quoted context omitted.
some other imporant quotes: "Average human off the street: 70-80%. STEM college grad: >95%. Panel of 10 random humans: 99-100%" -@fchollet on X So, considering that the $3400/task system isn't able to compete with STEM college grad yet, we still have some room (but it is shrinking, i expect even more compute will be thrown and we'll see these barriers broken in coming years) Also, some other back of envelope calculat…
I don't follow how 10 random humans can beat the average STEM college grad and average humans in that tweet. I suspect it's really "a panel of 10 randomly chosen experts in the space" or something? I agree the most interesting thing to watch will be cost for a given score more than maximum possible score achieved (not that the latter won't be interesting by any means).
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#527Just as an aside, I've personally found o1 to be completely useless for coding. Sonnet 3.5 remains the king of the hill by quite some margin
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#528Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#529Earlier quoted context omitted.
> ~=$3400 per single task report says it is $17 per task, and $6k for whole dataset of 400 tasks.
That's the low-compute mode. In the plot at the top where they score 88%, O3 High (tuned) is ~3.4k
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#530Congratulations to Francois Chollet on making the most interesting and challenging LLM benchmark so far. A lot of people have criticized ARC as not being relevant or indicative of true reasoning, but I think it was exactly the right thing. The fact that scaled reasoning models are finally showing progress on ARC proves that what it measures really is relevant and important for reasoning. It's obvious to everyone that…
There is a benchmark, NovelQA, that LLMs don't dominate when it feels like they should. The benchmark is to read a novel and answer questions about it. LLMs are below human evaluation, as I last looked, but it doesn't get much attention. Once it is passed, I'd like to see one that is solving the mystery in a mystery book right before it's revealed. We'd need unpublished mystery novels to use for that benchmark, but I…
I would think this is a not so good bench. Author does not write logically, they write for entertainment.