Live data from Hacker News

OpenAI O3 breakthrough high score on ARC-AGI-PUB

arcprize.org

961–970 of 1001 posts

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#962

Incredibly impressive. Still can't really shake the feeling that this is o3 gaming the system more than it is actually being able to reason. If the reasoning capabilities are there, there should be no reason why it achieves 90% on one version and 30% on the next. If a human maintains the same performance across the two versions, an AI with reason should too.

How would gaming the system work here? Is there some flaw in the way the tasks are generated?

[dead]

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#963

Let me go against some skeptics and explain why I think full o3 is pretty much AGI or at least embodies most essential aspects of AGI. What has been lacking so far in frontier LLMs is the ability to reliably deal with the right level of abstraction for a given problem. Reasoning is useful but often comes out lacking if one cannot reason at the right level of abstraction. (Note that many humans can't either when they…

on the spatial data i see it as a highly intelligent head of a machine that just needs better limbs and better senses. i think that's where most hardware startups will specialize with in the coming decades, different industries with different needs.

[deleted]

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#964
post #439

Earlier quoted context omitted.

In order to replace actual humans doing their job I think LLMs are lacking in judgement, sense of time and agenticism.

I mean fkcu me when they have those things, however, maybe they are just lazy and their judgement is fine, for a lazy intelligence. Inner-self thinks "why are these bastards asking me to do this? ". I doubt that is actually happening, but now, .. prove it isn't.

[deleted]

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#966

Earlier quoted context omitted.

Looks like quite shoddy code though. Like, the procedure for running a shell command is pure side-effect procedural code, neither returning the exit code of the command nor its output. Like the incomplete stackoverflow answer it probably was trained from. It might do one job at a time, but once this stuff gets integrated into one coherent thing, one needs to rewrite lots of it, to actually be composable. Though, of c…

Which code is shoddy? The Claude or o3-mini one? If you mean Claude, then have you checked the o3-mini one is better?

Youtube is currently blocking my VPN, can't watch it.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#967
post #738

How do the organisers keep the private test set private? Does openAI hand them the model for testing? If they use a model API, then surely OpenAI has access to the private test set questions and can include it in the next round of training? (I am sure I am missing something.)

Isn’t that why they call it “ Semi-Private”? There’s a fully private test set too as I understand it, that o3 hasn’t run on yet.

And o3 will not run on the private set unless it is a truly free and open source model (presumably also the case for ARC-AGI-2). This is the distinction between private and semi-private. In private you provide all the knowledge/weights/logic to operate without any external communication. Private benchmark results are the only true evaluation of performance on any benchmark -- reserved for a final evaluation. It is the only way to prevent shenanigans.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#968

Let me go against some skeptics and explain why I think full o3 is pretty much AGI or at least embodies most essential aspects of AGI. What has been lacking so far in frontier LLMs is the ability to reliably deal with the right level of abstraction for a given problem. Reasoning is useful but often comes out lacking if one cannot reason at the right level of abstraction. (Note that many humans can't either when they…

how about if it can work at a job? people can do that, can o3 do it?

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#969

Incredibly impressive. Still can't really shake the feeling that this is o3 gaming the system more than it is actually being able to reason. If the reasoning capabilities are there, there should be no reason why it achieves 90% on one version and 30% on the next. If a human maintains the same performance across the two versions, an AI with reason should too.

I think you've hit the nail on the head there. If these systems of reasoning are truly general then they should be able to perform consistently in the same way a human does across similar tasks, baring some variance.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#970

Terrifying. This news makes me happy I save all my money. My only hope for the future is that I can retire early before I’m unemployable

The whole economy is going to crash and money won't be worth anything, so it won't matter if you have money or not. Of course is a chance we will find ourselves in Utopia, but yeah, a chance.

Money buys real assets which will be worth something; AI can't magic up land or energy for instance. In fact AI is a dream for capital, and a nightmare for labor/work/human intelligence w.r.t value.
Post reply on HN