Live data from Hacker News

OpenAI O3 breakthrough high score on ARC-AGI-PUB

arcprize.org

391–400 of 1001 posts

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#391

Earlier quoted context omitted.

GPQA scores are mostly from pre-training, against content in the corpus. They have gone silent but look at the GPT4 technical report which calls this out. We are nowhere close to what Sam Altman calls AGI and transformers are still limited to what uniform-TC0 can do. As an example the Boolean Formula Value Problem is NC1-complete, thus beyond transformers but trivial to solve with a TM. As it is now proven that the f…

Isn't any physically realizable computer (including our brains) limited to what uniform-TC0 can do?

Do you just mean because any physically realizable computer is a finite state machine? Or...?

I wouldn't describe a computer's usual behavior as having constant depth.

It is fairly typical to talk about problems in P as being feasible (though when the constant factors are too big, this isn't strictly true of course).

Just because for unreasonably large inputs, my computer can't run a particular program and produce the correct answer for that input, due to my computer running out of memory, we don't generally say that my computer is fundamentally incapable of executing that algorithm.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#393

It sucks that I would love to be excited about this... but I mostly feel anxiety and sadness.

Same, it's sad but I honestly hoped they never achieved these results and it came out that it wasn't possible or would take an insurmountable amount of resources but here we are ok the verge of making most humans useless when it comes to productivity. While there are those that are excited, the world is not prepared for the level of distress this could put on the average person without critical changes at a monumenta…

If you don't feel like the world needed grand scale changes at a societal level with all the global problems we're unable to solve, you haven't been paying attention. Income inequality, corporate greed, political apathy, global warming.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#394

Efficiency is now key. ~=$3400 per single task to meet human performance on this benchmark is a lot. Also it shows the bullets as "ARC-AGI-TUNED", which makes me think they did some undisclosed amount of fine-tuning (eg. via the API they showed off last week), so even more compute went into this task. We can compare this roughly to a human doing ARC-AGI puzzles, where a human will take (high variance in my subjective…

some other imporant quotes: "Average human off the street: 70-80%. STEM college grad: >95%. Panel of 10 random humans: 99-100%" -@fchollet on X So, considering that the $3400/task system isn't able to compete with STEM college grad yet, we still have some room (but it is shrinking, i expect even more compute will be thrown and we'll see these barriers broken in coming years) Also, some other back of envelope calculat…

> are we stuck waiting for the 20-25 years for GPU improvements

If this turns out to be hard to optimize / improve then there will be a huge economic incentive for efficient ASICs. No freaking way we’ll be running on GPUs for 20-25 years, or even 2.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#395

Complete aside here: I used to do work with amputees and prosthetics. There is a standardized test (and I just cannot remember the name) that fits in a briefcase. It's used for measuring the level of damage to the upper limbs and for prosthetic grading. Basically, it's got the dumbest and simplest things in it. Stuff like a lock and key, a glass of water and jug, common units of currency, a zipper, etc. It tests if y…

I had a pretty bad case of tendinitis once, that basically made my thumb useless since using it would cause extreme pain. That test seems really good. I could use a computer keyboard without any issue, but putting a belt on or pouring water was impossible.

I had a swollen elbow a short while ago, and the amount of things I've never thought about that were affected by reduced elbow join mobility and an inability to put pressure on the elbow was disturbing.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#397
post #319

The cost to run the highest performance o3 model is estimated to be somewhere between $2,000 and $3,400 per task.[1] Based on these estimates, o3 costs about 100x what it would cost to have a human perform the exact same task. Many people are therefore dismissing the near-term impact of these models because of these extremely expensive costs. I think this is a mistake. Even if very high costs make o3 uneconomic for b…

Your economic analysis is deeply flawed. If there was anything that valuable and that required that much manpower, it would already have driven up the cost of labor accordingly. The one property that could conceivably justify a substantially higher cost is secrecy. After all, you can't (legally) kill a human after your project ends to ensure total secrecy. But that takes us into thriller novel territory.

I don't think that's right. Free societies don't tolerate total mobilization by their governments outside of war time, no matter how valuable the outcomes might be in the long term, in part because of the very economic impacts you describe. Human-level AI - even if it's very expensive - puts something that looks a lot like total mobilization within reach without the societal pushback. This is especially true when it comes to tasks that society as a whole may not sufficiently value, but that a state actor might value very much, and when paired with something like a co-located reactor and data center that does not impact the grid.

That said, this is all predicated on o3 or similar actually having achieved human level reasoning. That's yet to be fully proven. We'll see!

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#398

Complete aside here: I used to do work with amputees and prosthetics. There is a standardized test (and I just cannot remember the name) that fits in a briefcase. It's used for measuring the level of damage to the upper limbs and for prosthetic grading. Basically, it's got the dumbest and simplest things in it. Stuff like a lock and key, a glass of water and jug, common units of currency, a zipper, etc. It tests if y…

>We had hand prosthetics that could play Mozart at 5x speed on a baby grand

I'd love to know more about this.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#399
post #352

It is not exactly AGI but huge step toward it. I would expect this step in 2028-2030. I cant really understand why people are happy with it, this technology is so dangerous that can disrupt whole society. It's neither like smartphone nor internet. What will happen to 3rd world countries. Lots of unsolved questions and world is not prepared for such a change. Lots of people will lose their jobs I am not even mentionin…

> What will happen to 3rd world countries Probably less disruption than will happen in 1st world countries. > No one will have chance to be rich anymore It's strange to reach this conclusion from "look, a massive new productivity increase".

Intelligence is the thing distinguishing humans from all previous inventions that already were superhuman in some narrow domain.

car : horse :: AGI : humans

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#400

It is not exactly AGI but huge step toward it. I would expect this step in 2028-2030. I cant really understand why people are happy with it, this technology is so dangerous that can disrupt whole society. It's neither like smartphone nor internet. What will happen to 3rd world countries. Lots of unsolved questions and world is not prepared for such a change. Lots of people will lose their jobs I am not even mentionin…

Same, I don’t really get the excitement. None of these companies are pushing for a utopian Star Trek society either with that power.

Open models will catch up next year or the year after, there only so many things to try and there's lots of people trying them, so it's more or less an inevitability.

The part to get excited about is that there's plenty of headroom left to gain in performance. They called o1 a preview, and it was, a preview for QwQ and similar models. We get the demo from OAI and then get the real thing for free next year.

Post reply on HN