Very cool. I recommend scrolling down to look at the example problem that O3 still can’t solve. It’s clear what goes on in the human brain to solve this problem: we look at one example, hypothesize a simple rule that explains it, and then check that hypothesis against the other examples. It doesn’t quite work, so we zoom into an example that we got wrong and refine the hypothesis so that it solves that sample. We kee…
I took a look at those examples that o3 can't solve. Looks similar to an IQ-test. Took me less time to figure out the 3 examples that it took to read your post. I was honestly a bit surprised to see how visual the tasks were. I had thought they were text based. So now I'm quite impressed that o3 can solve this type of task at all.
OpenAI O3 breakthrough high score on ARC-AGI-PUB
341–350 of 1001 posts
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#342Based on the chart, the Kaggle SOTA model is far more impressive. These O3 models are more expensive to run than just hiring a mechanical turk worker. It's nice we are proving out the scaling hypothesis further, it's just grossly inelegant. The Kaggle SOTA performs 2x as well as o1 high at a fraction of the cost
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#343Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#344Congratulations to Francois Chollet on making the most interesting and challenging LLM benchmark so far. A lot of people have criticized ARC as not being relevant or indicative of true reasoning, but I think it was exactly the right thing. The fact that scaled reasoning models are finally showing progress on ARC proves that what it measures really is relevant and important for reasoning. It's obvious to everyone that…
This benchmark has done a wonderful job with marketing by picking a great name. It's largely irrelevant for LLMs despite the fact it's difficult.
Consider how much of the model is just noise for a task like this given the low amount of information in each token and the high embedding dimensions used in LLMs.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#345I feel like AI is already changing how we work and live - I've been using it myself for a lot of my development work. Though, what I'm really concerned about is what happens when it gets smart enough to do pretty much everything better (or even close) than humans can. We're talking about a huge shift where first knowledge workers get automated, then physical work too. The thing is, our whole society is built around p…
I'll get concerned when it stops sucking so hard. It's like talking to a dumb robot. Which it unsurprisingly is.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#346Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#347Congratulations to Francois Chollet on making the most interesting and challenging LLM benchmark so far. A lot of people have criticized ARC as not being relevant or indicative of true reasoning, but I think it was exactly the right thing. The fact that scaled reasoning models are finally showing progress on ARC proves that what it measures really is relevant and important for reasoning. It's obvious to everyone that…
This emphasizes persons and a self-conceived victory narrative over the ground truth. Models have regularly made progress on it, this is not new with the o-series. Doing astoundingly well on it, and having a mutually shared PR interest with OpenAI in this instance, doesn't mean a pile of visual puzzles is actually AGI or some well thought out and designed benchmark of True Intelligence(tm). It's one type of visual pu…
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#348Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#349Congratulations to Francois Chollet on making the most interesting and challenging LLM benchmark so far. A lot of people have criticized ARC as not being relevant or indicative of true reasoning, but I think it was exactly the right thing. The fact that scaled reasoning models are finally showing progress on ARC proves that what it measures really is relevant and important for reasoning. It's obvious to everyone that…
"The fact that scaled reasoning models are finally showing progress on ARC proves that what it measures really is relevant and important for reasoning." Not sure I understand how this follows. The fact that a certain type of model does well on a certain benchmark means that the benchmark is relevant for a real-world reasoning? That doesn't make sense.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#350Earlier quoted context omitted.
> ~=$3400 per single task report says it is $17 per task, and $6k for whole dataset of 400 tasks.
That's the low-compute mode. In the plot at the top where they score 88%, O3 High (tuned) is ~3.4k
Sorry for being thick Im just confused how they can turn this into an addordable service?