Earlier quoted context omitted.
Well, it would be AGI if you could connect a camera to it to solve it, similar to how blind people would be able to solve it if you restored their eyesight. But if the lack of vision is a fundamental limitation of their architecture, then it seems more fair not to call them AGI.
People blind from birth literally lack the neural circuits to comprehend visual data. Are they not intelligent?
ARC-AGI-3
251–260 of 394 posts
Re: ARC-AGI-3
#252Earlier quoted context omitted.
So what? Are you suggesting that an agent exhibiting genuine AGI will be tripped up by having to ingest json rather than rgb pixels? LLMs are largely trained on textual data so json is going to be much closer to whatever native is for them. But by all means, give the agents access to an API that returns pixel data. However I fully expect that would reduce performance rather than increase it.
Because it is. Opus 4.6 jumps from 0.0% to 97.1% when given visual input
However, if it can't figure out to render the json to a visual on its own does it really qualify as AGI? I'd still say the benchmark is doing its job here. Granted it's not a perfectly even playing field in that case but I think the goal is to test for progress towards AGI as opposed to hosting a fair tournament.
Re: ARC-AGI-3
#253Earlier quoted context omitted.
If you only outdo humans 50% of the time you're never going to get consensus on if you've qualified. Whereas outdoing 90% of humans on 90% of all the most difficult tasks we could come up with is going to be difficult to argue against. This benchmark is only one such task. After this one there's still the rest of that 90% to go. Beating humans isn't anywhere near sufficient to qualify as ASI. That's an entirely diffe…
Even dumb humans are considered to have general intelligence. If the bar is having to outdo the median human, then 50% of humans don't have general intelligence.
Frontier models are reliably providing high undergraduate to low graduate level customized explanations of highly technical topics at this point. Yet I regularly catch them making errors that a human never would and which betray a fatal lack of any sort of mental model. What are we supposed to make of that?
It's an exceedingly weird situation we find ourselves in. These models can provide useful assistance to literal mathematicians yet simultaneously show clear evidence of lacking some sort of reasoning the details of which I find difficult to articulate. They also can't learn on the job whatsoever. Is that intelligence? Probably. But is it general? I don't think so, at least not in the sense that "AGI" implies to me.
Once humanity runs out of examples that reliably trip them up I'll agree that they're "general" to the same extent that humans are regardless of if we've figured out the secrets behind things such as cohesive world models, self awareness, active learning during operation, and theory of mind.
Re: ARC-AGI-3
#254Earlier quoted context omitted.
Because it is. Opus 4.6 jumps from 0.0% to 97.1% when given visual input
That's impressive. I'm also a bit surprised - I wouldn't have expected it to be trained much at all on that sort of visual input task. I think I'd be similarly surprised to learn that a frontier model was particularly good at playing retro videogames or actuating a robot for example. However, if it can't figure out to render the json to a visual on its own does it really qualify as AGI? I'd still say the benchmark is…
Can you render serialized JSON text blob to a visual with your brain only? The model can't do anything better than this - no harness means no tool at all, no way to e.g. implement a visualizer in whatever programming language and run it.
Why don't human testers receive the same JSON text blob and no visualizers? It's like giving human testers a harness (a playable visualizer), but deliberately cripples it for the model.
Re: ARC-AGI-3
#255Earlier quoted context omitted.
I’d be hesitant to call that ASI if it’s pretty obvious how you’d write a regular old program to solve it.
It’s not that simple since each problem is supposed to be distinct and different enough that no single program can solve multiple of them properly. No problem spec is provided as well iiuc so you can’t simply ask an LLM to generate code without doing other things.
Re: ARC-AGI-3
#256Earlier quoted context omitted.
Isn’t this what AGI is by design? People CAN learn to become good at videogames. Modern LLMs can’t, they have to be retrained from scratch (I consider pre-training to be a completely different process than learning). I also don’t necessarily agree that a grandma would fail. Give her enough motivation and a couple days and she’ll manage these. My main criticism would be that it doesn’t seem like this test allows onlin…
What I'm saying is that this test is just another "out-of-distribution task" for an LLM. And it will be solved using the exact same methods we always use: it will end up in the pre-training data, and LLMs will crush it. This has absolutely nothing to do with AGI. Once they beat these tests, new ones will pop up. They'll beat those, and people will invent the next batch. The way I see it, the true formula for AGI is:…
that's exactly the point! once we cannot invent the next batch (that is easy for humans to solve), that will be AGI
Re: ARC-AGI-3
#257Earlier quoted context omitted.
That's impressive. I'm also a bit surprised - I wouldn't have expected it to be trained much at all on that sort of visual input task. I think I'd be similarly surprised to learn that a frontier model was particularly good at playing retro videogames or actuating a robot for example. However, if it can't figure out to render the json to a visual on its own does it really qualify as AGI? I'd still say the benchmark is…
> However, if it can't figure out to render the json to a visual on its own does it really qualify as AGI? I'd still say the benchmark is doing its job here. Can you render serialized JSON text blob to a visual with your brain only? The model can't do anything better than this - no harness means no tool at all, no way to e.g. implement a visualizer in whatever programming language and run it. Why don't human testers…
Re: ARC-AGI-3
#258Earlier quoted context omitted.
There's a very simple solution to this problem here. Instead of wink-wink-nudge-nudge implying that 100% is 'human baseline', calculate the median human score from the data you already have and put it on that chart.
Its below 1% lmao
Re: ARC-AGI-3
#259Earlier quoted context omitted.
How do you know they aren't conscious of we don't know what consciousness is, and have no test to see if anyone or anything is conscious? This may seem like a joke, but your answer will likely be in the vain of "conscious things are obviously conscious", which gets us nowhere. I mean, self motivation and a desire to not be turned off can be programmed into even decades old AIs.
As a philosophical zombie myself[0], I'm well aware of how hard it is to define and test consciousness. That's why I tried to clarify what I meant with: desire for self-preservation and intrinsic motivation. Which LLMs clearly lack, don't you agree? Also, I'm not saying that those things couldn't be programmed in, just that so far, they don't seem necessary . [0] I lack a conscious experience and qualia
Re: ARC-AGI-3
#260Earlier quoted context omitted.
> As long as there is a gap between AI and human learning, we do not have AGI. Don't read the statement as a human dunk on LLMs, or even as philosophy. The gap is important because of its special and devastating economic consequences. When the gap becomes truly zero, all human knowledge work is replaceable. From there, with robots, its a short step to all work is replaceable. What's worse, the condition is sufficient…
The gap is important because of its special and devastating economic consequences. When the gap becomes truly zero, all human knowledge work is replaceable. From there, with robots, its a short step to all work is replaceable. I don’t know why statements like this are just taken as gospel fact. There are plenty of economic activities which do not disappear even if an AI can do them. Here’s one: I support certain arti…
Yeah, but obviously no human can clear that bar either.
> Here’s another: positions that have deep experience in certain industries and have valuable networks
What stops an AGI from gaining "deep experience in an industry"? Or forming networks? There's plenty of popular bot accounts across social media already.