Live data from Hacker News

OpenAI O3 breakthrough high score on ARC-AGI-PUB

arcprize.org

321–330 of 1001 posts

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#321

It sucks that I would love to be excited about this... but I mostly feel anxiety and sadness.

Same, it's sad but I honestly hoped they never achieved these results and it came out that it wasn't possible or would take an insurmountable amount of resources but here we are ok the verge of making most humans useless when it comes to productivity.

While there are those that are excited, the world is not prepared for the level of distress this could put on the average person without critical changes at a monumental level.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#322

Many are incorrectly citing 85% as human-level performance. 85% is just the (semi-arbitrary) threshold for the winning the prize. o3 actually beats the human average by a wide margin: 64.2% for humans vs. 82.8%+ for o3. ... Here's the full breakdown by dataset, since none of the articles make it clear -- Private Eval: - 85%: threshold for winning the prize [1] Semi-Private Eval: - 87.5%: o3 (unlimited compute) [2] -…

[deleted]

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#323

Very cool. I recommend scrolling down to look at the example problem that O3 still can’t solve. It’s clear what goes on in the human brain to solve this problem: we look at one example, hypothesize a simple rule that explains it, and then check that hypothesis against the other examples. It doesn’t quite work, so we zoom into an example that we got wrong and refine the hypothesis so that it solves that sample. We kee…

I took a look at those examples that o3 can't solve. Looks similar to an IQ-test.

Took me less time to figure out the 3 examples that it took to read your post.

I was honestly a bit surprised to see how visual the tasks were. I had thought they were text based. So now I'm quite impressed that o3 can solve this type of task at all.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#324

Congratulations to Francois Chollet on making the most interesting and challenging LLM benchmark so far. A lot of people have criticized ARC as not being relevant or indicative of true reasoning, but I think it was exactly the right thing. The fact that scaled reasoning models are finally showing progress on ARC proves that what it measures really is relevant and important for reasoning. It's obvious to everyone that…

> making the most interesting and challenging LLM benchmark so far. This[1] is currently the most challenging benchmark. I would like to see how O3 handles it, as O1 solved only 1%. 1. https://epoch.ai/frontiermath/the-benchmark

You're right, I was wrong to say "most challenging" as there have been harder ones coming out recently. I think the correct statement would be "most challenging long-standing benchmark" as I don't believe any other test designed in 2019 has resisted progress for so long. FrontierMath is only a month old. And of course the real key feature of ARC is that it is easy for humans. FrontierMath is (intentionally) not.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#325

Many are incorrectly citing 85% as human-level performance. 85% is just the (semi-arbitrary) threshold for the winning the prize. o3 actually beats the human average by a wide margin: 64.2% for humans vs. 82.8%+ for o3. ... Here's the full breakdown by dataset, since none of the articles make it clear -- Private Eval: - 85%: threshold for winning the prize [1] Semi-Private Eval: - 87.5%: o3 (unlimited compute) [2] -…

If my life depended on the average rando solving 8/10 arc-prize puzzles, I'd consider myself dead.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#326

It sucks that I would love to be excited about this... but I mostly feel anxiety and sadness.

Humanity is about to enter an even steeper hockey stick growth curve. Progressing along the Kardashev scale feels all but inevitable. We will live to see Longevity Escape Velocity. I'm fucking pumped and feel thrilled and excited and proud of our species. Sure, there will be growing pains, friction, etc. Who cares? There always is with world-changing tech. Always.

> Sure, there will be growing pains, friction, etc. Who cares?

That's right. Who cares about pains of others and why they even should are absolutely words to live by.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#327
post #221

This might sound dumb, and I'm not sure how to phrase this, but is there a way to measure the raw model output quality without all the more "traditional" engineering work (mountain of `if` statements I assume) done on top of the output? And if so, would that be a better measure of when scaling up the input data will start showing diminishing returns? (I know very little about the guts of LLMs or how they're tested, s…

what do you mean by the mountain of if-statements on top of the output? like checking if the output matches the expected result in evaluations?

Like when you type something into the chat gpt app I am guessing it will start by preprocessing your input, doing some sanity checks, making sure it doesn’t say “how do I build a bomb?” or whatever. It may or may not alter/clean up your input before sending it to the model for processing. Once processed, there’s probably dozens of services it goes through to detect if the output is racist, somehow actually contained a bomb recipe, or maybe copywriter material, normal pattern matching stuff, maybe some advanced stuff like sentiment analysis to see if the output is bad mouthing Trump or something, and it might either alter the output or simply try again.

I’m wondering when you strip out all that “extra” non-model pre and post processing, if there’s someway to measure performance of that.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#328

Complete aside here: I used to do work with amputees and prosthetics. There is a standardized test (and I just cannot remember the name) that fits in a briefcase. It's used for measuring the level of damage to the upper limbs and for prosthetic grading. Basically, it's got the dumbest and simplest things in it. Stuff like a lock and key, a glass of water and jug, common units of currency, a zipper, etc. It tests if y…

I think assembling Legos would be a cool robot benchmark: you need to parse the instructions, locate the pieces you need, pick them up, orient them, snap them to your current assembly, visually check if you achieved the desired state, repeat

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#329
Intelligence comes in many forms and flavors. ARC prize questions are just one version of it -- perhaps measuring more human-like pattern recognition than true intelligence.

Can machines be more human-like in their pattern recognition? O3 met this need today.

While this is some form of accomplishment, it's nowhere near the scientific and engineering problem solving needed to call something truly artificial (human-like) intelligent.

What’s exciting is that these reasoning models are making significant strides in tackling eng and scientific problem-solving. Solving the ARC challenge seems almost trivial in comparison to that.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#330

It sucks that I would love to be excited about this... but I mostly feel anxiety and sadness.

Humanity is about to enter an even steeper hockey stick growth curve. Progressing along the Kardashev scale feels all but inevitable. We will live to see Longevity Escape Velocity. I'm fucking pumped and feel thrilled and excited and proud of our species. Sure, there will be growing pains, friction, etc. Who cares? There always is with world-changing tech. Always.

I would rather follow in the steps of uncle Ted than let AI turn me in a homeless person. It’s no consolation that my tent will have a nice view of a lunar colony
Post reply on HN