Live data from Hacker News

OpenAI O3 breakthrough high score on ARC-AGI-PUB

arcprize.org

521–530 of 1001 posts

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#521

With only a 100x increase in cost, we improved performance by 0.1x and continued plotting this concave-down diminishing-returns type graph! Hurray for logarithmic x-axes! Joking aside, better than ever before at any cost is an achievement, it just doesn't exactly scream "breakthrough" to me.

[flagged]

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#522

With only a 100x increase in cost, we improved performance by 0.1x and continued plotting this concave-down diminishing-returns type graph! Hurray for logarithmic x-axes! Joking aside, better than ever before at any cost is an achievement, it just doesn't exactly scream "breakthrough" to me.

It may eventually be able to solve any problem

Ah. Me, too.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#523
post #521

With only a 100x increase in cost, we improved performance by 0.1x and continued plotting this concave-down diminishing-returns type graph! Hurray for logarithmic x-axes! Joking aside, better than ever before at any cost is an achievement, it just doesn't exactly scream "breakthrough" to me.

[flagged]

Agi level my ass

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#524

With only a 100x increase in cost, we improved performance by 0.1x and continued plotting this concave-down diminishing-returns type graph! Hurray for logarithmic x-axes! Joking aside, better than ever before at any cost is an achievement, it just doesn't exactly scream "breakthrough" to me.

imo it's a mistake to interpret the marginal increases in the upper echelons of benchmarks as materially marginal gains. Chess is an example. ELO narrows heavily at the top, but each ELO point carries more relative weight. This is a bit apples and oranges since chess is adversarial, but I think the point stands.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#525

Direct quote from the ARC-AGI blog: “SO IS IT AGI? ARC-AGI serves as a critical benchmark for detecting such breakthroughs, highlighting generalization power in a way that saturated or less demanding benchmarks cannot. However, it is important to note that ARC-AGI is not an acid test for AGI – as we've repeated dozens of times this year. It's a research tool designed to focus attention on the most challenging unsolve…

> acid test The css acid test? This can be gamed too.

https://en.wikipedia.org/wiki/Acid_test:

> An acid test is a qualitative chemical or metallurgical assay utilizing acid. Historically, it often involved the use of a robust acid to distinguish gold from base metals. Figuratively, the term represents any definitive test for attributes, such as gauging a person's character or evaluating a product's performance.

Specifically here, they're using the figurative sense of "definitive test".

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#526

Earlier quoted context omitted.

some other imporant quotes: "Average human off the street: 70-80%. STEM college grad: >95%. Panel of 10 random humans: 99-100%" -@fchollet on X So, considering that the $3400/task system isn't able to compete with STEM college grad yet, we still have some room (but it is shrinking, i expect even more compute will be thrown and we'll see these barriers broken in coming years) Also, some other back of envelope calculat…

I don't follow how 10 random humans can beat the average STEM college grad and average humans in that tweet. I suspect it's really "a panel of 10 randomly chosen experts in the space" or something? I agree the most interesting thing to watch will be cost for a given score more than maximum possible score achieved (not that the latter won't be interesting by any means).

ARC-AGI is essentially an IQ test. There is no "expert in the space". Its just a question of if youre able to spot the pattern.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#529

Earlier quoted context omitted.

> ~=$3400 per single task report says it is $17 per task, and $6k for whole dataset of 400 tasks.

That's the low-compute mode. In the plot at the top where they score 88%, O3 High (tuned) is ~3.4k

The low compute one did as well as the average person though

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#530

Congratulations to Francois Chollet on making the most interesting and challenging LLM benchmark so far. A lot of people have criticized ARC as not being relevant or indicative of true reasoning, but I think it was exactly the right thing. The fact that scaled reasoning models are finally showing progress on ARC proves that what it measures really is relevant and important for reasoning. It's obvious to everyone that…

There is a benchmark, NovelQA, that LLMs don't dominate when it feels like they should. The benchmark is to read a novel and answer questions about it. LLMs are below human evaluation, as I last looked, but it doesn't get much attention. Once it is passed, I'd like to see one that is solving the mystery in a mystery book right before it's revealed. We'd need unpublished mystery novels to use for that benchmark, but I…

> I'd like to see one that is solving the mystery in a mystery book right before it's revealed.

I would think this is a not so good bench. Author does not write logically, they write for entertainment.

Post reply on HN