As an aside, I'm a little miffed that the benchmark calls out "AGI" in the name, but then heavily cautions that it's necessary but insufficient for AGI. > ARC-AGI serves as a critical benchmark for detecting such breakthroughs, highlighting generalization power in a way that saturated or less demanding benchmarks cannot. However, it is important to note that ARC-AGI is not an acid test for AGI
OpenAI O3 breakthrough high score on ARC-AGI-PUB
421–430 of 1001 posts
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#422Earlier quoted context omitted.
These comments are getting ridiculous. I remember when this test was first discussed here on HN and everyone agreed that it clearly proves current AI models are not "intelligent" (whatever that means). And people tried to talk me down when I theorised this test will get nuked soon - like all the ones before. It's time people woke up and realised that the old age of AI is over. This new kind is here to stay and it wil…
> It's time people woke up and realised that the old age of AI is over. This new kind is here to stay and it will take over the world. And you better guess it'll be sooner rather than later and start to prepare. I was just thinking about how 3D game engines were perceived in the 90s. Every six months some new engine came out, blew people's minds, was declared photorealistic, and was forgotten a year later. The best o…
The transition seems to map well to the point where engines got sophisticated enough, that highly dedicated high-schoolers couldn't keep up. Until then, people would routinely make hobby game engines (for games they'd then never finish) that were MVPs of what the game industry had a year or three earlier. I.e. close enough to compete on visuals with top photorealistic games of a given year - but more importantly, this was a time where you could do cool nerdy shit to impress your friends and community.
Then Unreal and Unity came out, with a business model that killed the motivation to write your own engine from scratch (except for purely educational purposes), we got more games, more progress, but the excitement was gone.
Maybe it's just a spurious correlation, but it seems to track with:
> and only became so again when OpenAI's InstructGPT was demonstrated.
Which is again, if you exclude training SOTA models - which is still mostly out of reach for anyone but a few entities on the planet - the time where anyone can do something cool that doesn't have a better market alternative yet, and any dedicated high-schooler can make truly impressive and useful work, outpacing commercial and academic work based on pure motivation and focus alone (it's easier when you're not being distracted by bullshit incentives like user growth or making VCs happy or churning out publications, farming citations).
It's, once again, a time of dreams, where anyone with some technical interest and a bit of free time can make the future happen in front of their eyes.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#423Complete aside here: I used to do work with amputees and prosthetics. There is a standardized test (and I just cannot remember the name) that fits in a briefcase. It's used for measuring the level of damage to the upper limbs and for prosthetic grading. Basically, it's got the dumbest and simplest things in it. Stuff like a lock and key, a glass of water and jug, common units of currency, a zipper, etc. It tests if y…
> We had hand prosthetics that could play Mozart at 5x speed on a baby grand, but could not pick up a silver dollar or zip a jacket even a little bit. " I must be missing something, how can they be able to play Mozart at 5x speed with their prosthetics but not zip a jacket? They could press keys but not do tasks requiring feedback? Or did you mean they used to play Mozart at 5x speed before they became amputees?
My point is -- being able to zip a jacket is all about those subtle actions, and could actually be harder than "just" playing piano fast.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#424The programming task they gave o3-mini high (creating Python server that allows chatting with OpenAI API and run some code in terminal) didn't seem very hard? Strange choice of example for something that's claimed to be a big step forwards. YT timestamped link: https://www.youtube.com/watch?v=SKBG1sqdyIU&t=768s (thanks for the fixed link @photonboom) Updated: I gave the task to Claude 3.5 Sonnet and it worked first s…
I've been doing similar stuff in Claude for months and it's not that impressive when you see how limited they really are when going non boilerplate.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#425OpenAI spent approximately $1,503,077 to smash the SOTA on ARC-AGI with their new o3 model semi-private evals (100 tasks): 75.7% @ $2,012 total/100 tasks (~$20/task) with just 6 samples & 33M tokens processed in ~1.3 min/task and a cost of $2012 The “low-efficiency” setting with 1024 samples scored 87.5% but required 172x more compute. If we assume compute spent and cost are proportional, then OpenAI might have just…
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#426Interesting that in the video, there is an admission that they have been targeting this benchmark. A comment that was quickly shut down by Sam. A bit puzzling to me. Why does it matter ?
In reality it seems to be a bit of both - there is some general intelligence based on having been "trained on the internet", but it seems these super-human math/etc skills are very much from them having focused on training on those.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#427How can there be "private" taks when you have use the OpenAI API to run queries? OpenAI sees everything.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#428O3 High (tuned) model scored an 88% at what looks like $6,000/task haha I think soon we'll be pricing any kind of tasks by their compute costs. So basically, human = $50/task, AI = $6,000/task, use human. If AI beats human, use AI? Ofc that's considering both get 100% scores on the task
Compute costs on AI with the same roughly the same capabilities have been halving every ~7 months. That makes something like this competitive in ~3 years
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#429Seriously, programming as a profession will end soon. Let's not kid us anymore. Time to jump the ship.
Why specifically programming? I think every knowledge profession is at risk, or at the very minimum suspect to a huge transformation. Doctors, analysts, lawyers, etc.