Live data from Hacker News

ARC-AGI-3

arcprize.org

251–260 of 394 posts

Re: ARC-AGI-3

#251

Earlier quoted context omitted.

Well, it would be AGI if you could connect a camera to it to solve it, similar to how blind people would be able to solve it if you restored their eyesight. But if the lack of vision is a fundamental limitation of their architecture, then it seems more fair not to call them AGI.

People blind from birth literally lack the neural circuits to comprehend visual data. Are they not intelligent?

I think they don't actually lack them, or lack only a small fraction (their brains are ≈99% like a normal human brain), such that if they were an AI model, they could be fairly trivially upgraded with vision capability.

Re: ARC-AGI-3

#252

Earlier quoted context omitted.

So what? Are you suggesting that an agent exhibiting genuine AGI will be tripped up by having to ingest json rather than rgb pixels? LLMs are largely trained on textual data so json is going to be much closer to whatever native is for them. But by all means, give the agents access to an API that returns pixel data. However I fully expect that would reduce performance rather than increase it.

Because it is. Opus 4.6 jumps from 0.0% to 97.1% when given visual input

That's impressive. I'm also a bit surprised - I wouldn't have expected it to be trained much at all on that sort of visual input task. I think I'd be similarly surprised to learn that a frontier model was particularly good at playing retro videogames or actuating a robot for example.

However, if it can't figure out to render the json to a visual on its own does it really qualify as AGI? I'd still say the benchmark is doing its job here. Granted it's not a perfectly even playing field in that case but I think the goal is to test for progress towards AGI as opposed to hosting a fair tournament.

Re: ARC-AGI-3

#253

Earlier quoted context omitted.

If you only outdo humans 50% of the time you're never going to get consensus on if you've qualified. Whereas outdoing 90% of humans on 90% of all the most difficult tasks we could come up with is going to be difficult to argue against. This benchmark is only one such task. After this one there's still the rest of that 90% to go. Beating humans isn't anywhere near sufficient to qualify as ASI. That's an entirely diffe…

Even dumb humans are considered to have general intelligence. If the bar is having to outdo the median human, then 50% of humans don't have general intelligence.

Not true. We don't have a good definition for intelligence - it's very much an I'll know it when I see it sort of thing.

Frontier models are reliably providing high undergraduate to low graduate level customized explanations of highly technical topics at this point. Yet I regularly catch them making errors that a human never would and which betray a fatal lack of any sort of mental model. What are we supposed to make of that?

It's an exceedingly weird situation we find ourselves in. These models can provide useful assistance to literal mathematicians yet simultaneously show clear evidence of lacking some sort of reasoning the details of which I find difficult to articulate. They also can't learn on the job whatsoever. Is that intelligence? Probably. But is it general? I don't think so, at least not in the sense that "AGI" implies to me.

Once humanity runs out of examples that reliably trip them up I'll agree that they're "general" to the same extent that humans are regardless of if we've figured out the secrets behind things such as cohesive world models, self awareness, active learning during operation, and theory of mind.

Re: ARC-AGI-3

#254

Earlier quoted context omitted.

Because it is. Opus 4.6 jumps from 0.0% to 97.1% when given visual input

That's impressive. I'm also a bit surprised - I wouldn't have expected it to be trained much at all on that sort of visual input task. I think I'd be similarly surprised to learn that a frontier model was particularly good at playing retro videogames or actuating a robot for example. However, if it can't figure out to render the json to a visual on its own does it really qualify as AGI? I'd still say the benchmark is…

> However, if it can't figure out to render the json to a visual on its own does it really qualify as AGI? I'd still say the benchmark is doing its job here.

Can you render serialized JSON text blob to a visual with your brain only? The model can't do anything better than this - no harness means no tool at all, no way to e.g. implement a visualizer in whatever programming language and run it.

Why don't human testers receive the same JSON text blob and no visualizers? It's like giving human testers a harness (a playable visualizer), but deliberately cripples it for the model.

Re: ARC-AGI-3

#255
post #170

Earlier quoted context omitted.

I’d be hesitant to call that ASI if it’s pretty obvious how you’d write a regular old program to solve it.

It’s not that simple since each problem is supposed to be distinct and different enough that no single program can solve multiple of them properly. No problem spec is provided as well iiuc so you can’t simply ask an LLM to generate code without doing other things.

A human can sit down to play a game with unknown rules and write a spec as he goes. If a model can't even figure out to attempt that, let alone succeed at it, then it most certainly isn't an example of "general" intelligence.

Re: ARC-AGI-3

#256
post #200

Earlier quoted context omitted.

Isn’t this what AGI is by design? People CAN learn to become good at videogames. Modern LLMs can’t, they have to be retrained from scratch (I consider pre-training to be a completely different process than learning). I also don’t necessarily agree that a grandma would fail. Give her enough motivation and a couple days and she’ll manage these. My main criticism would be that it doesn’t seem like this test allows onlin…

What I'm saying is that this test is just another "out-of-distribution task" for an LLM. And it will be solved using the exact same methods we always use: it will end up in the pre-training data, and LLMs will crush it. This has absolutely nothing to do with AGI. Once they beat these tests, new ones will pop up. They'll beat those, and people will invent the next batch. The way I see it, the true formula for AGI is:…

> . Once they beat these tests, new ones will pop up. They'll beat those, and people will invent the next batch.

that's exactly the point! once we cannot invent the next batch (that is easy for humans to solve), that will be AGI

Re: ARC-AGI-3

#257
post #254

Earlier quoted context omitted.

That's impressive. I'm also a bit surprised - I wouldn't have expected it to be trained much at all on that sort of visual input task. I think I'd be similarly surprised to learn that a frontier model was particularly good at playing retro videogames or actuating a robot for example. However, if it can't figure out to render the json to a visual on its own does it really qualify as AGI? I'd still say the benchmark is…

> However, if it can't figure out to render the json to a visual on its own does it really qualify as AGI? I'd still say the benchmark is doing its job here. Can you render serialized JSON text blob to a visual with your brain only? The model can't do anything better than this - no harness means no tool at all, no way to e.g. implement a visualizer in whatever programming language and run it. Why don't human testers…

Huh. I thought it wasn't supposed to receive any instructions tailored to the task but I didn't understand it to be restricted from accessing truly general tools such as programming languages. To do otherwise is to require pointless hoop jumping as frontier models inevitably get retrained to play games using a json (or other arbitrary) representation at which point it will be natural for them and the real test will begin.

Re: ARC-AGI-3

#258

Earlier quoted context omitted.

There's a very simple solution to this problem here. Instead of wink-wink-nudge-nudge implying that 100% is 'human baseline', calculate the median human score from the data you already have and put it on that chart.

Its below 1% lmao

where did you get this 1%?

Re: ARC-AGI-3

#259

Earlier quoted context omitted.

How do you know they aren't conscious of we don't know what consciousness is, and have no test to see if anyone or anything is conscious? This may seem like a joke, but your answer will likely be in the vain of "conscious things are obviously conscious", which gets us nowhere. I mean, self motivation and a desire to not be turned off can be programmed into even decades old AIs.

As a philosophical zombie myself[0], I'm well aware of how hard it is to define and test consciousness. That's why I tried to clarify what I meant with: desire for self-preservation and intrinsic motivation. Which LLMs clearly lack, don't you agree? Also, I'm not saying that those things couldn't be programmed in, just that so far, they don't seem necessary . [0] I lack a conscious experience and qualia

How can you tell that you lack conscious experience and qualia?

Re: ARC-AGI-3

#260
post #139

Earlier quoted context omitted.

> As long as there is a gap between AI and human learning, we do not have AGI. Don't read the statement as a human dunk on LLMs, or even as philosophy. The gap is important because of its special and devastating economic consequences. When the gap becomes truly zero, all human knowledge work is replaceable. From there, with robots, its a short step to all work is replaceable. What's worse, the condition is sufficient…

The gap is important because of its special and devastating economic consequences. When the gap becomes truly zero, all human knowledge work is replaceable. From there, with robots, its a short step to all work is replaceable. I don’t know why statements like this are just taken as gospel fact. There are plenty of economic activities which do not disappear even if an AI can do them. Here’s one: I support certain arti…

> AI can only use data it has access to, and it’s never going to have access to everyone’s individual brain everywhere at all times.

Yeah, but obviously no human can clear that bar either.

> Here’s another: positions that have deep experience in certain industries and have valuable networks

What stops an AGI from gaining "deep experience in an industry"? Or forming networks? There's plenty of popular bot accounts across social media already.

Post reply on HN