OpenAI O3 breakthrough high score on ARC-AGI-PUB
861–870 of 1001 posts
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#862My initial impression: it's very impressive and very exciting. My skeptical impression: it's complete hubris to conflate ARC or any benchmark with truly general intelligence. I know my skepticism here is identical to moving goalposts. More and more I am shifting my personal understanding of general intelligence as a phenomenon we will only ever be able to identify with the benefit of substantial retrospect. As it is…
I think it's still an interesting way to measure general intellience, it's just that o3 has demonstrated that you can actually achieve human performance on it by training it on the public training set and giving it ridiculous amounts of compute, which I imagine equates to ludicrously long chains-of-thought, and if I understand correctly more than one chain-of-thought per task (they mention sample sizes in the blog po…
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#863I just graduated college, and this was a major blow. I studied Mechanical Engineering and went into Sales Engineering because cause I love technology and people, but articles like this do nothing but make me dread the future. I have no idea what to specialize in, what skills I should master, or where I should be spending my time to build a successful career. Seems like we’re headed toward a world where you automate s…
This is me as well. Either: 1) Just give up computing entirely, the field I've been dreaming about since childhood. Perhaps if I immiserate myself with a dry regulated engineering field or trade I would perhaps survive to recursive self-improvement, but if anything the length it takes to pivot (I am a Junior in College that has already done probably 3/4th of my CS credits) means I probably couldn't get any foothold u…
It's a massive bubble, and things like these "benchmarks" are all part of the hype game. Is the tech cool and useful? For sure, but anyone trying to tell you this benchmark is in any way proof of AGI and will replace everyone is either an idiot or more likely has a vested interest in you believing them. OpenAI's whole marketing shtick is to scare people into thinking their next model is "too dangerous" to be released thus driving up hype, only to release it anyway and for it to fall flat on its face.
Also, if there's any jobs LLMs can replace right now, it's the useless managerial and C-suite, not the people doing the actual work. If these people weren't charlatans they'd be the first ones to go while pushing this on everyone else.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#864I guess I get to brag now. ARC AGI has no real defences against Big Data, memorisation-based approaches like LLMs. I told you so: https://news.ycombinator.com/item?id=42344336 And that answers my question about fchollet's assurances that LLMs without TTT (Test Time Training) can't beat ARC AGI: [me] I haven't had the chance to read the papers carefully. Have they done ablation studies? For instance, is the following…
How are the Bongard Problems going?
Interestingly, Bongard problems do not have a private test set, unlike ARC-AGI. Can that be because they don't need it? Is it possible that Bongard Problems are a true test of (visual) reasoning that requires intelligence to be solved?
Ooooh! Frisson of excitement!
But I guess it's just that nobody remembers them and so nobody has seriously tried to solve them with Big Data stuff.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#865Earlier quoted context omitted.
> making the most interesting and challenging LLM benchmark so far. This[1] is currently the most challenging benchmark. I would like to see how O3 handles it, as O1 solved only 1%. 1. https://epoch.ai/frontiermath/the-benchmark
You're right, I was wrong to say "most challenging" as there have been harder ones coming out recently. I think the correct statement would be "most challenging long-standing benchmark" as I don't believe any other test designed in 2019 has resisted progress for so long. FrontierMath is only a month old. And of course the real key feature of ARC is that it is easy for humans. FrontierMath is (intentionally) not.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#866Earlier quoted context omitted.
some other imporant quotes: "Average human off the street: 70-80%. STEM college grad: >95%. Panel of 10 random humans: 99-100%" -@fchollet on X So, considering that the $3400/task system isn't able to compete with STEM college grad yet, we still have some room (but it is shrinking, i expect even more compute will be thrown and we'll see these barriers broken in coming years) Also, some other back of envelope calculat…
It's also worth keeping in mind that AIs are a lot less risky to deploy for businesses than humans. You can scale them up and down at any time, they can work 24/7 (including holidays) with no overtime pay and no breaks, they need no corporate campuses, office space, HR personnel or travel budgets, you don't have to worry about key employees going on sick/maternity leave or taking time off the moment they're needed mo…
For AI example(s): Attribution is low, a system built without human intervention may suddenly fall outside its own expertise and hallucinate itself into a corner, everyone may just throw more compute at a system until it grows without bound, etc etc.
This "You can scale up to infinity" problem might become "You have to scale up to infinity" to build any reasonably sized system with AI. The shovel-sellers get fantastically rich but the businesses are effectively left holding the risk from a fast-moving, unintuitive, uninspected, partially verified codebase. I just don't see how anyone not building a CRUD app/frontend could be comfortable with that, but then again my Tesla is effectively running such a system to drive me and my kids. Albeit, that's on a well-defined problem and within literally human-made guardrails.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#867Someone asked if true intelligence requires a foundation of prior knowledge. This is the way I think about it. I = E / K where I is the intelligence of the system, E is the effectiveness of the system, and K is the prior knowledge. For example, a math problem is given to two students, each solving the problem with the same effectiveness (both get the correct answer in the same amount of time). However, student A happ…
There should be also a factor about resource consumption. See here: https://lorenzopieri.com/pgii/
Yes, resource consumption is important. But your car guzzling a lot of gas doesn't mean he drives slower. It just means it drives slower per mol of petrol consumed.
It's good to know whether your system has a high or low 'bang for buck' metric, but that doesn't directly affect how much bang you get.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#868Efficiency is now key. ~=$3400 per single task to meet human performance on this benchmark is a lot. Also it shows the bullets as "ARC-AGI-TUNED", which makes me think they did some undisclosed amount of fine-tuning (eg. via the API they showed off last week), so even more compute went into this task. We can compare this roughly to a human doing ARC-AGI puzzles, where a human will take (high variance in my subjective…
Fundamentally it's a search through some enormous state space. Advancements are "tricks" that let us find useful subsets more efficiently.
Zooming way out, we have a bunch of social tricks, hardware tricks, and algorithmic tricks that have resulted in a super useful subset. It's not the subset that we want though, so the hunt continues.
Hopefully it doesn't require revising too much in the hardware & social bag of tricks, those are lot more painful to revisit...
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#869Earlier quoted context omitted.
> ~=$3400 per single task report says it is $17 per task, and $6k for whole dataset of 400 tasks.
You're misreading it, there's two different runs, a low and a high compute run. The number for the high-compute one is ~172x the first one according to the article so ~=$2900
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#870Isn’t this like a brute force approach? Given it costs $ 3000 per task, thats like 600 GPU hours (h100 at Azure) In that amount of time the model can generate millions of chains of thoughts and then spend hours reviewing them or even testing them out one by one. Kind of like trying until something sticks and that happens to solve 80% of ARC. I feel like reasoning works differently in my brain. ;)