Can I just say what a dick move it was to do this as a 12 days of Christmas. I mean to be honest I agree with the arguments this isn’t as impressive as my initial impression, but they clearly intended it to be shocking/a show of possible AGI, which is rightly scary. It feels so insensitive to that right before a major holiday when the likely outcome is a lot of people feeling less secure in their career/job/life. Tha…
The vast majority of people who will lose jobs to AI aren’t following AGI benchmarks, or even know what AGI is short for.
OpenAI O3 breakthrough high score on ARC-AGI-PUB
831–840 of 1001 posts
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#832Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#833Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#834Efficiency is now key. ~=$3400 per single task to meet human performance on this benchmark is a lot. Also it shows the bullets as "ARC-AGI-TUNED", which makes me think they did some undisclosed amount of fine-tuning (eg. via the API they showed off last week), so even more compute went into this task. We can compare this roughly to a human doing ARC-AGI puzzles, where a human will take (high variance in my subjective…
> Super exciting that OpenAI pushed the compute out this far it's even more exciting than that. the fact that you even can use more compute to get more intelligence is a breakthrough. if they spent even more on inference, would they get even better scores on arc agi?
I'm not so sure—what they're doing by just throwing more tokens at it is similar to "solving" the traveling salesman problem by just throwing tons of compute into a breadth first search. Sure, you can get better and better answers the more compute you throw at it (with diminishing returns), but is that really that surprising to anyone who's been following tree of thought models?
All it really seems to tell us is that the type of model that OpenAI has available is capable of solving many of the types of problems that ARC-AGI-PUB has set up given enough compute time. It says nothing about "intelligence" as the concept exists in most people's heads—it just means that a certain very artificial (and intentionally easy for humans) class of problem that wasn't computable is now computable if you're willing to pay an enormous sum to do it. A breakthrough of sorts, sure, but not a surprising one given what we've seen already.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#835How do the organisers keep the private test set private? Does openAI hand them the model for testing? If they use a model API, then surely OpenAI has access to the private test set questions and can include it in the next round of training? (I am sure I am missing something.)
If we really want to imagine a cold-war-style solution, the two teams could meet in an empty warehouse, bring one computer with the model, one with the benchmarks, and connect them with a USB cable. In practice I assume they just gave them the benchmarks and took it on the honor system they wouldn't cheat, yeah. They can always cook up a new test set for next time, it's only 10% of the benchmark content anyway and th…
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#836Earlier quoted context omitted.
Still it's comparing average human level performance with best AI performance. Examples of things o3 failed at are insanely easy for humans.
There are things Chimps do easily that humans fail at, and vice/versa of course. There are blind spots, doesn't take away from 'general'.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#837How do the organisers keep the private test set private? Does openAI hand them the model for testing? If they use a model API, then surely OpenAI has access to the private test set questions and can include it in the next round of training? (I am sure I am missing something.)
There’s a fully private test set too as I understand it, that o3 hasn’t run on yet.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#838Efficiency is now key. ~=$3400 per single task to meet human performance on this benchmark is a lot. Also it shows the bullets as "ARC-AGI-TUNED", which makes me think they did some undisclosed amount of fine-tuning (eg. via the API they showed off last week), so even more compute went into this task. We can compare this roughly to a human doing ARC-AGI puzzles, where a human will take (high variance in my subjective…
some other imporant quotes: "Average human off the street: 70-80%. STEM college grad: >95%. Panel of 10 random humans: 99-100%" -@fchollet on X So, considering that the $3400/task system isn't able to compete with STEM college grad yet, we still have some room (but it is shrinking, i expect even more compute will be thrown and we'll see these barriers broken in coming years) Also, some other back of envelope calculat…
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#839Seriously, programming as a profession will end soon. Let's not kid us anymore. Time to jump the ship.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#840Earlier quoted context omitted.
Still it's comparing average human level performance with best AI performance. Examples of things o3 failed at are insanely easy for humans.
There are things Chimps do easily that humans fail at, and vice/versa of course. There are blind spots, doesn't take away from 'general'.