Live data from Hacker News

OpenAI O3 breakthrough high score on ARC-AGI-PUB

arcprize.org

81–90 of 1001 posts

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#81
post #19

My initial impression: it's very impressive and very exciting. My skeptical impression: it's complete hubris to conflate ARC or any benchmark with truly general intelligence. I know my skepticism here is identical to moving goalposts. More and more I am shifting my personal understanding of general intelligence as a phenomenon we will only ever be able to identify with the benefit of substantial retrospect. As it is…

" it's complete hubris to conflate ARC or any benchmark with truly general intelligence."

Maybe it would help to include some human results in the AI ranking.

I think we'd find that Humans score lower?

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#82
post #73

Great results. However, let's all just admit it. It has well replaced journalists, artists and on its way to replace nearly both junior and senior engineers. The ultimate intention of "AGI" is that it is going to replace tens of millions of jobs. That is it and you know it. It will only accelerate and we need to stop pretending and coping. Instead lets discuss solutions for those lost jobs. So what is the replacement…

Do you follow Jack Clark? I noticed he's been on the road a lot talking to governments and policy makers, and not just in the "AI is coming" way he used to talk.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#83
post #24

O3 High (tuned) model scored an 88% at what looks like $6,000/task haha I think soon we'll be pricing any kind of tasks by their compute costs. So basically, human = $50/task, AI = $6,000/task, use human. If AI beats human, use AI? Ofc that's considering both get 100% scores on the task

Compute costs on AI with the same roughly the same capabilities have been halving every ~7 months.

That makes something like this competitive in ~3 years

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#84
post #19

My initial impression: it's very impressive and very exciting. My skeptical impression: it's complete hubris to conflate ARC or any benchmark with truly general intelligence. I know my skepticism here is identical to moving goalposts. More and more I am shifting my personal understanding of general intelligence as a phenomenon we will only ever be able to identify with the benefit of substantial retrospect. As it is…

From the statement where - this is a pretty tough test where AI scores low vs humans just last year, and AI can do it as good as humans may not be AGI which I agree, but it means something with all caps

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#85
post #40
post #24

O3 High (tuned) model scored an 88% at what looks like $6,000/task haha I think soon we'll be pricing any kind of tasks by their compute costs. So basically, human = $50/task, AI = $6,000/task, use human. If AI beats human, use AI? Ofc that's considering both get 100% scores on the task

Isn't that generally what ... all jobs are? Automation Cost vs Longterm Human cost... its why amazon did the weird "our stores are AI driven" but in reality was cheaper to higher a bunch of guys in a sweat shop to look at the cameras and write things down lol. The thing is given what we've seen from distillation and tech, even if its 6,000/task... that will come down drastically over time through optimization and jus…

I remember hearing Tesla trying to automate all of production but some things just couldn’t , like the wiring which humans still had to do.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#86

Just as an aside, I've personally found o1 to be completely useless for coding. Sonnet 3.5 remains the king of the hill by quite some margin

To fill this out, I find o1-pro (and -preview when it was live) to be pretty good at filling in blindspots/spotting holistic bugs. I use Claude for day to day, and when Claude is spinning, o1 often can point out why. It's too slow for AI coding, and I agree that at default its responses aren't always satisfying.

That said, I think its code style is arguably better, more concise and has better patterns -- Claude needs a fair amount of prompting and oversight to not put out semi-shitty code in terms of structure and architecture.

In my mind: going from Slowest to Fastest, and Best Holistically to Worst, the list is:

1. o1-pro 2. Claude 3.5 3. Gemini 2 Flash

Flash is so fast, that it's tempting to use more, but it really needs to be kept to specific work on strong codebases without complex interactions.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#87
post #19

My initial impression: it's very impressive and very exciting. My skeptical impression: it's complete hubris to conflate ARC or any benchmark with truly general intelligence. I know my skepticism here is identical to moving goalposts. More and more I am shifting my personal understanding of general intelligence as a phenomenon we will only ever be able to identify with the benefit of substantial retrospect. As it is…

> My skeptical impression: it's complete hubris to conflate ARC or any benchmark with truly general intelligence.

But isn’t it interesting to have several benchmarks? Even if it’s not about passing the Turing test, benchmarks serve a purpose—similar to how we measure microprocessors or other devices. Intelligence may be more elusive, but even if we had an oracle delivering the ultimate intelligence benchmark, we'd still argue about its limitations. Perhaps we'd claim it doesn't measure creativity well, and we'd find ourselves revisiting the same debates about different kinds of intelligences.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#88
post #52

Whenever a benchmark that was thought to be extremely difficult is (nearly) solved, it's a mix of two causes. One is that progress on AI capabilities was faster than we expected, and the other is that there was an approach that made the task easier than we expected. I feel like the there's a lot of the former here, but the compute cost per task (thousands of dollars to solve one little color grid puzzle??) suggests t…

> the other is that there was an approach that made the task easier than we expected.

from reading Dennett's philosophy, I'm convinced that that's how human intelligence works - for each task that "only a human could do that", there's a trick that makes it easier than it seems. We are bags of tricks.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#90

So now not only are the models closed, but so are their evals?! This is a "semi-private" eval. WTH is that supposed to mean? I'm sure the model is great but I refuse to take their word for it.

The private evaluation set is private from the public/OpenAI so companies can't train on those problems and cheat their way to a high score by overfitting.

If the models run on OpenAIs servers then surely they could still see the questions being put into it if they wanted to cheat? That could only be prevented by making the evaluation a one-time deal that can't be repeated, or by having OpenAI distribute their models for evaluators to run themselves, which I doubt they're inclined to do.
Post reply on HN