Uhhhh… It was trained on ARC data? So they targeted a specific benchmark and are surprised and blown away the LLM performed well in it? What’s that law again? When a benchmark is targeted by some system the benchmark becomes useless?
OpenAI O3 breakthrough high score on ARC-AGI-PUB
491–500 of 1001 posts
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#492I would like to see this repeated with my highly innovative HARC-HAGI, which is ARC-AGI but it uses hexagons instead of squares. I suspect humans would only make slightly more brain farts on HARC-HAGI than ARC-AGI, but O3 would fail very badly since it almost certainly has been specifically trained on squares. I am not really trying to downplay O3. But this would be a simple test as to whether O3 is truly "a system c…
Taking this a level of abstraction higher, I expect that in the next couple of years we'll see systems like o3 given a runtime budget that they can use for training/fine-tuning smaller models in an ad-hoc manner.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#493Earlier quoted context omitted.
some other imporant quotes: "Average human off the street: 70-80%. STEM college grad: >95%. Panel of 10 random humans: 99-100%" -@fchollet on X So, considering that the $3400/task system isn't able to compete with STEM college grad yet, we still have some room (but it is shrinking, i expect even more compute will be thrown and we'll see these barriers broken in coming years) Also, some other back of envelope calculat…
Let's say that Google is already 1 generation ahead of nvidia in terms of efficient AI compute. ($1700) Then let's say that OpenAI brute forced this without any meta-optimization of the hypothesized search component (they just set a compute budget). This is probably low hanging fruit and another 2x in compute reduction. ($850) Then let's say that OpenAI was pushing really really hard for the numbers and was willing t…
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#494Congratulations to Francois Chollet on making the most interesting and challenging LLM benchmark so far. A lot of people have criticized ARC as not being relevant or indicative of true reasoning, but I think it was exactly the right thing. The fact that scaled reasoning models are finally showing progress on ARC proves that what it measures really is relevant and important for reasoning. It's obvious to everyone that…
"The fact that scaled reasoning models are finally showing progress on ARC proves that what it measures really is relevant and important for reasoning." Not sure I understand how this follows. The fact that a certain type of model does well on a certain benchmark means that the benchmark is relevant for a real-world reasoning? That doesn't make sense.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#495Let me go against some skeptics and explain why I think full o3 is pretty much AGI or at least embodies most essential aspects of AGI. What has been lacking so far in frontier LLMs is the ability to reliably deal with the right level of abstraction for a given problem. Reasoning is useful but often comes out lacking if one cannot reason at the right level of abstraction. (Note that many humans can't either when they…
Have we really watered down the definition of AGI that much? LLMs aren't really capable of "learning" anything outside their training data. Which I feel is a very basic and fundamental capability of humans. Every new request thread is a blank slate utilizing whatever context you provide for the specific task and after the tread is done (or context limit runs out) it's like it never happened. Sure you can use database…
ChatGPT has had for some time the feature of storing memories about its conversations with users. And you can use function calling to make this more generic.
I think drawing the boundary at “model + scaffolding” is more interesting.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#496Let me go against some skeptics and explain why I think full o3 is pretty much AGI or at least embodies most essential aspects of AGI. What has been lacking so far in frontier LLMs is the ability to reliably deal with the right level of abstraction for a given problem. Reasoning is useful but often comes out lacking if one cannot reason at the right level of abstraction. (Note that many humans can't either when they…
In order to replace actual humans doing their job I think LLMs are lacking in judgement, sense of time and agenticism.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#497If we're inferring the answers of the block patterns from minimal or no additional training, it's very impressive, but how much time have they had to work on O3 after sharing puzzle data with O1? Seems there's some room for questionable antics!
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#498Earlier quoted context omitted.
some other imporant quotes: "Average human off the street: 70-80%. STEM college grad: >95%. Panel of 10 random humans: 99-100%" -@fchollet on X So, considering that the $3400/task system isn't able to compete with STEM college grad yet, we still have some room (but it is shrinking, i expect even more compute will be thrown and we'll see these barriers broken in coming years) Also, some other back of envelope calculat…
I don't follow how 10 random humans can beat the average STEM college grad and average humans in that tweet. I suspect it's really "a panel of 10 randomly chosen experts in the space" or something? I agree the most interesting thing to watch will be cost for a given score more than maximum possible score achieved (not that the latter won't be interesting by any means).
This isn't to say groups always outperform their members on all tasks, just that it isn't unusual to see a result like that.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#499This is also wildly ahead in SWE-bench (71.7%, previous 48%) and Frontier Math (25% on high compute, previous 2%). So much for a plateau lol.
You're talking apples and oranges. The plateau the frontier models have hit is the limited further gains to be had from dataset (+ corresponding model/compute) scaling. These new reasoning models are taking things in a new direction basically by adding search (inference time compute) on top of the basic LLM. So, the capabilities of the models are still improving, but the new variable is how deep of a search you want…
They found a way to make test time compute a lot more effective and that is an advance but the idea is not new, the architecture is not new.
And the vast majority of people convinced LLMs plateaued did so regardless of test time compute.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#500This is so impressive that it brings out the pessimist in me. Hopefully my skepticism will end up being unwarranted, but how confident are we that the queries are not routed to human workers behind the API? This sounds crazy but is plausible for the fake-it-till-you-make-it crowd. Also given the prohibitive compute costs per task, typical users won't be using this model, so the scheme could go on for quite sometime b…
this is an impressive tinfoil take. but what would be their plan in the medium term? like once they release this people can check their data
In the medium term the plan could be to achieve AGI, and then AGI would figure out how to actually write o3. (Probably after AGI figures out the business model though: https://www.reddit.com/r/MachineLearning/s/OV4S2hGgW8)