Live data from Hacker News

OpenAI O3 breakthrough high score on ARC-AGI-PUB

arcprize.org

531–540 of 1001 posts

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#531
post #521

With only a 100x increase in cost, we improved performance by 0.1x and continued plotting this concave-down diminishing-returns type graph! Hurray for logarithmic x-axes! Joking aside, better than ever before at any cost is an achievement, it just doesn't exactly scream "breakthrough" to me.

[flagged]

I know AGI is a bit of a moving goalpost these days, but by my personally anecdotal and irrelevant opinion this ain’t AGI.

I’ll let those smarter than me debate the merits of AGI, but if it can’t learn and self-improve it isn’t “general” intelligence.

This is a very smart computer, accomplishing a very niche set of problems. Cool? Yes. AGI? No.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#532

Efficiency is now key. ~=$3400 per single task to meet human performance on this benchmark is a lot. Also it shows the bullets as "ARC-AGI-TUNED", which makes me think they did some undisclosed amount of fine-tuning (eg. via the API they showed off last week), so even more compute went into this task. We can compare this roughly to a human doing ARC-AGI puzzles, where a human will take (high variance in my subjective…

> I am interpreting this result as human level reasoning now costs (approximately) 41k/hr to 2.5M/hr with current compute.

On a very simple, toy task, which arc-agi basically is. Arc-agi tests are not hard per se, just LLM’s find them hard. We do not know how this scales for more complex, real world tasks.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#533
post #342
post #166

Based on the chart, the Kaggle SOTA model is far more impressive. These O3 models are more expensive to run than just hiring a mechanical turk worker. It's nice we are proving out the scaling hypothesis further, it's just grossly inelegant. The Kaggle SOTA performs 2x as well as o1 high at a fraction of the cost

But does that Kaggle solution achieve human level perf with any level of compute? I think you're missing the forest for the trees here.

The article says the ensemble of Kaggle solutions (aggregated in some unexplained way) achieves 81%. This is better than their average Mechanical Turk worker, but worse than their average STEM grad. It's better than tuned o3 with low compute, worse than tuned o3 with high compute.

There's also a point on the figure marked "Kaggle SOTA", around 60%. I can't find any explanation for that, but I guess it's the best individual Kaggle solution.

The Kaggle solutions would probably score higher with more compute, but nobody has any incentive to spend >$1M on approaches that obviously don't generalize. OpenAI did have this incentive to spend tuning and testing o3, since it's possible that will generalize to a practically useful domain (but not yet demonstrated). Even if it ultimately doesn't, they're getting spectacular publicity now from that promise.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#534

Earlier quoted context omitted.

Highly challenging for LLMs because it has nothing to do with language. LLMs and their training processes have all kinds of optimizations for language and how it's presented. This benchmark has done a wonderful job with marketing by picking a great name. It's largely irrelevant for LLMs despite the fact it's difficult. Consider how much of the model is just noise for a task like this given the low amount of informati…

The benchmark is designed to test for AGI and intelligence, specifically the ability to solve novel problems. If the hypothesis is that LLMs are the “computer” that drives the AGI then of course the benchmark is relevant in testing for AGI. I don’t think you understand the benchmark and its motivation. ARC AGI benchmark problems are extremely easy and simple for humans. But LLMs fail spectacularly at them. Why they f…

> The benchmark is designed to test for AGI and intelligence, specifically the ability to solve novel problems.

It's a bunch of visual puzzles. They aren't a test for AGI because it's not general. If models (or any other system for that matter) could solve it, we'd be saying "this is a stupid puzzle, it has no practical significance". It's a test of some sort of specific intelligence. On top of that, the vast majority of blind people would fail - are they not generally intelligent?

The name is marketing hype.

The benchmark could be called "random puzzles LLMs are not good at because they haven't been optimized for it because it's not valuable benchmark". Sure, it wasn't designed for LLMs, but throwing LLMs at it and saying "see?" is dumb. We can throw in benchmarks for tennis playing, chess playing, video game playing, car driving and a bajillion other things while we are at it.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#535

Earlier quoted context omitted.

There is a benchmark, NovelQA, that LLMs don't dominate when it feels like they should. The benchmark is to read a novel and answer questions about it. LLMs are below human evaluation, as I last looked, but it doesn't get much attention. Once it is passed, I'd like to see one that is solving the mystery in a mystery book right before it's revealed. We'd need unpublished mystery novels to use for that benchmark, but I…

> I'd like to see one that is solving the mystery in a mystery book right before it's revealed. I would think this is a not so good bench. Author does not write logically, they write for entertainment.

So I'm thinking of something like Locked-room mystery where the idea is it's solvable, and the reader is given a chance to solve.

The reason it seems like an interesting bench, is it's a puzzle presented in a long context. Its like testing if an LLm is at Sherlock Holmes level of world and motivation modelling.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#536

With only a 100x increase in cost, we improved performance by 0.1x and continued plotting this concave-down diminishing-returns type graph! Hurray for logarithmic x-axes! Joking aside, better than ever before at any cost is an achievement, it just doesn't exactly scream "breakthrough" to me.

compute gets cheaper and cheaper every year. This model will be in your phone by 2030 if we continue at the pace we've been at the last few years.

There’s probably enough VC money to subsidize the costs for a few more years.

But the data centres running the training for models like this are bringing up new methane power plants at a fast rate at a time when we need to be reducing reliance on O&G.

But let’s assume that the efficiency gains out pace the resource consumption with the help of all the subsidies being thrown in and we achieve AGI.

What’s the benefit? Do we get more fresh water?

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#537
I just graduated college, and this was a major blow. I studied Mechanical Engineering and went into Sales Engineering because cause I love technology and people, but articles like this do nothing but make me dread the future.

I have no idea what to specialize in, what skills I should master, or where I should be spending my time to build a successful career.

Seems like we’re headed toward a world where you automate someone else’s job or be automated yourself.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#538

I'm 22 and have no clue what I'm meant to do in a world where this is a thing. I'm moving to a semi rural, outdoorsy area where they teach data science and marine science and I can enjoy my days hiking, and the march of technology is a little slower. I know this will disrupt so much of our way of life, so I'm chasing what fun innocent years are left before things change dramatically.

If information technology workers become twice as productive, you’ll want more of them for your business, not less.

There are way more data analysts now than when it required paper and pencil.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#540
post #246

OpenAI spent approximately $1,503,077 to smash the SOTA on ARC-AGI with their new o3 model semi-private evals (100 tasks): 75.7% @ $2,012 total/100 tasks (~$20/task) with just 6 samples & 33M tokens processed in ~1.3 min/task and a cost of $2012 The “low-efficiency” setting with 1024 samples scored 87.5% but required 172x more compute. If we assume compute spent and cost are proportional, then OpenAI might have just…

Pretty sure this "cost" is based on their retail price instead of actual inference cost.

Yeah and can run off peak, etc.

Does seem to show an absolutely massive market for inference compute…

Post reply on HN