Live data from Hacker News

OpenAI O3 breakthrough high score on ARC-AGI-PUB

arcprize.org

341–350 of 1001 posts

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#341

Very cool. I recommend scrolling down to look at the example problem that O3 still can’t solve. It’s clear what goes on in the human brain to solve this problem: we look at one example, hypothesize a simple rule that explains it, and then check that hypothesis against the other examples. It doesn’t quite work, so we zoom into an example that we got wrong and refine the hypothesis so that it solves that sample. We kee…

I took a look at those examples that o3 can't solve. Looks similar to an IQ-test. Took me less time to figure out the 3 examples that it took to read your post. I was honestly a bit surprised to see how visual the tasks were. I had thought they were text based. So now I'm quite impressed that o3 can solve this type of task at all.

I also took some time to look at the ones it couldn't solve. I stopped after this one: https://kts.github.io/arc-viewer/page6/#47996f11

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#342
post #166

Based on the chart, the Kaggle SOTA model is far more impressive. These O3 models are more expensive to run than just hiring a mechanical turk worker. It's nice we are proving out the scaling hypothesis further, it's just grossly inelegant. The Kaggle SOTA performs 2x as well as o1 high at a fraction of the cost

But does that Kaggle solution achieve human level perf with any level of compute? I think you're missing the forest for the trees here.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#344

Congratulations to Francois Chollet on making the most interesting and challenging LLM benchmark so far. A lot of people have criticized ARC as not being relevant or indicative of true reasoning, but I think it was exactly the right thing. The fact that scaled reasoning models are finally showing progress on ARC proves that what it measures really is relevant and important for reasoning. It's obvious to everyone that…

Highly challenging for LLMs because it has nothing to do with language. LLMs and their training processes have all kinds of optimizations for language and how it's presented.

This benchmark has done a wonderful job with marketing by picking a great name. It's largely irrelevant for LLMs despite the fact it's difficult.

Consider how much of the model is just noise for a task like this given the low amount of information in each token and the high embedding dimensions used in LLMs.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#345

I feel like AI is already changing how we work and live - I've been using it myself for a lot of my development work. Though, what I'm really concerned about is what happens when it gets smart enough to do pretty much everything better (or even close) than humans can. We're talking about a huge shift where first knowledge workers get automated, then physical work too. The thing is, our whole society is built around p…

> Though, what I'm really concerned about is what happens when it gets smart enough to do pretty much everything better (or even close)

I'll get concerned when it stops sucking so hard. It's like talking to a dumb robot. Which it unsurprisingly is.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#347

Congratulations to Francois Chollet on making the most interesting and challenging LLM benchmark so far. A lot of people have criticized ARC as not being relevant or indicative of true reasoning, but I think it was exactly the right thing. The fact that scaled reasoning models are finally showing progress on ARC proves that what it measures really is relevant and important for reasoning. It's obvious to everyone that…

This emphasizes persons and a self-conceived victory narrative over the ground truth. Models have regularly made progress on it, this is not new with the o-series. Doing astoundingly well on it, and having a mutually shared PR interest with OpenAI in this instance, doesn't mean a pile of visual puzzles is actually AGI or some well thought out and designed benchmark of True Intelligence(tm). It's one type of visual pu…

100%. The hype is misguided. I doubt half the people excited about the result have even looked at what the benchmark is.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#349

Congratulations to Francois Chollet on making the most interesting and challenging LLM benchmark so far. A lot of people have criticized ARC as not being relevant or indicative of true reasoning, but I think it was exactly the right thing. The fact that scaled reasoning models are finally showing progress on ARC proves that what it measures really is relevant and important for reasoning. It's obvious to everyone that…

"The fact that scaled reasoning models are finally showing progress on ARC proves that what it measures really is relevant and important for reasoning." Not sure I understand how this follows. The fact that a certain type of model does well on a certain benchmark means that the benchmark is relevant for a real-world reasoning? That doesn't make sense.

It shows objectively that the models are getting better at some form of reasoning, which is at least worth noting. Whether that improved reasoning is relevant for the real world is a different question.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#350

Earlier quoted context omitted.

> ~=$3400 per single task report says it is $17 per task, and $6k for whole dataset of 400 tasks.

That's the low-compute mode. In the plot at the top where they score 88%, O3 High (tuned) is ~3.4k

sorry to be a noob, but can someone tell me doe sths mena o3 will be unaffordable for a typical user? Will only companies with thousands to spend per query be able to use this?

Sorry for being thick Im just confused how they can turn this into an addordable service?

Post reply on HN