Live data from Hacker News

ARC-AGI-3

arcprize.org

371–380 of 394 posts

Re: ARC-AGI-3

#371
post #200

Earlier quoted context omitted.

Isn’t this what AGI is by design? People CAN learn to become good at videogames. Modern LLMs can’t, they have to be retrained from scratch (I consider pre-training to be a completely different process than learning). I also don’t necessarily agree that a grandma would fail. Give her enough motivation and a couple days and she’ll manage these. My main criticism would be that it doesn’t seem like this test allows onlin…

What I'm saying is that this test is just another "out-of-distribution task" for an LLM. And it will be solved using the exact same methods we always use: it will end up in the pre-training data, and LLMs will crush it. This has absolutely nothing to do with AGI. Once they beat these tests, new ones will pop up. They'll beat those, and people will invent the next batch. The way I see it, the true formula for AGI is:…

good post, but I disagree Surival Function is needed for AGI. Why do you think Survival Function is needed?

The item I think you should add is a Mesolimbic System (Reward / Motivation). I think AGI needs motivation to direct its learning and tasks.

Also, I don't think the industry has just been training LLMs with more data to get advancement the last 2 years. RAG / Agents loops / skills / context mgmt are all just early forms a Memory system. An LLM with an updatable working set memory is a lot more capable than just an LLM.

Re: ARC-AGI-3

#372

Earlier quoted context omitted.

So is a person suffering from amnesia conscious if they lack short-term and long-term memory? Ruling out consciousness or qualia emerging from the inference in an LLM is just as invalid of a take as being 100% certain of its consciousness. We don’t know what consciousness really is, so only thing we can say with certainty is we do not know.

No, by continuity I mean literally moment to moment. Sorry if I didn’t clarify that. Even people with amnesia are still present moment to moment. As far as I know there are no things that we call conscious which have zero continuity. I think consciousness is not an abstract property in the world, therefore it’s tied to certain types of entities. Therefore an AI is not going to be “conscious” in the way an animal is,…

[dead]

Re: ARC-AGI-3

#373

Earlier quoted context omitted.

It has nothing to do with circular reasoning or my personal opinions. You can choose to define general intelligence in a way that excludes regular people if you like, but then you'd be using a weird definition that differs from how 99.9% of people define it. Humans have general intelligence by any common definition.

Defining it that way doesn't exclude ordinary people. That's an erroneous claim on your part. Humans as a class exhibit certain capabilities. Thus we expect a class of algorithm to either roughly meet or exceed those capabilities across the board in order to be considered "general". It is clear that we have not yet achieved that.

First, what is your definition exactly? That it must be better than the median human intelligence?

You're trying to define a term in a way that's completely detached from how anyone uses it. If we discover an alien race with an IQ of 95, people aren't going to say they don't have general intelligence.

We haven't defined an exact cutoff for what counts as general intelligence, but it has to include regular people with an IQ in the 70s that don't have a serious mental disability. If an AI can do every single cognitive task as well as a stupid person, it would have to qualify as having general intelligence if the stupid person qualified. It doesn't matter if the AI beats the median person 0% of the time, as long as it beats someone who is considered to have general intelligence at the task.

Re: ARC-AGI-3

#374

Earlier quoted context omitted.

1) Pointing out what tools to use is part of the intelligence that LLMs aren't great at. 2) one of the tools is a path finding algorithm. A big improvement/crutch over a regular LLM that has no such capability. You'd think if LLMs are intelligent they'd be able to determine that a path finding algorithm is necessary and have a sub agent code it up real quick. But apparently they just can't do that without humans step…

>You'd think if LLMs are intelligent they'd be able to determine that a path finding algorithm is necessary and have a sub agent code it up real quick. ARC 3 doesn't allow that so. >Here's the paper on what they did for the Duke Harness: https://blog.alexisfox.dev/arcagi3 Yeah, and the tools are general, not 'baked into the harness by the humans who coded it for this specific challenge.'

Adding a path finding algorithm and environment transform tools to a supposed "AGI", sure does seem like cheating to me. Sad part is, it's a cheat that only works on environments where pathfinding is a major part. And when it doesn't have those clues it bombs on everything.

I guess you really want to love the current SOTA LLMs. It's a shame they're dumb af.

Have a great day.

Re: ARC-AGI-3

#376

Earlier quoted context omitted.

>You'd think if LLMs are intelligent they'd be able to determine that a path finding algorithm is necessary and have a sub agent code it up real quick. ARC 3 doesn't allow that so. >Here's the paper on what they did for the Duke Harness: https://blog.alexisfox.dev/arcagi3 Yeah, and the tools are general, not 'baked into the harness by the humans who coded it for this specific challenge.'

Adding a path finding algorithm and environment transform tools to a supposed "AGI", sure does seem like cheating to me. Sad part is, it's a cheat that only works on environments where pathfinding is a major part. And when it doesn't have those clues it bombs on everything. I guess you really want to love the current SOTA LLMs. It's a shame they're dumb af. Have a great day.

>Adding a path finding algorithm and environment transform tools to a supposed "AGI", sure does seem like cheating to me.

You would need all that if you, a human wanted any chance of solving this benchmark in the format LLMs are given. The funny thing about this benchmark is that we don't even know how solvable it is, because the baseline is tested with radically different inputs.

>I guess you really want to love the current SOTA LLMs. It's a shame they're dumb af.

I guess you really don't want to think critically. Yeah good day lol.

Re: ARC-AGI-3

#377
post #54

Earlier quoted context omitted.

> AGI is usually defined as the ability to do any intellectual task about as well as a highly competent human could I think one major disconnect, is that for most people, AGI is when interacting with an AI is basically in every way like interacting with a human, including in failure modes. And likely, that this human would be the smartest most knowledgeable human you can imagine, like the top expert in all domains, w…

By that definition, does a human at the other end of a high-latency video call not have AGI because they can't react any faster that the connection's latency would allow them to have? From your POV what's the difference between that and an AI that's just slow?

> does a human at the other end of a high-latency video call not have AGI because they can't react any faster that the connection's latency would allow them to have

Correct. A person who'd mentally operate that slowly would be considered to have some cognitive disability. For example, would likely not be allowed to drive a car.

You could be fooled in thinking it is a human behind a slow connection, but layman would not consider it real AGI in my opinion, since you have to handicap the human, it seems like lowering the bar just to pretend you reached AGI.

You might recognize it's pretty close to AGI, if it has all the other qualities, but it needs to also operate at a similar response time, uptime, and so on.

My point is, everyone that's not trying to build AGI defines it as, same as an idealized smartest human would be in every way. I truly think this is how most people imagine AGI in their head, and until you have that, they'll say it's not AGI, and industry folks will claim the goalpost keeps moving, when in reality they kept setting their own post.

Re: ARC-AGI-3

#378

Earlier quoted context omitted.

Adding a path finding algorithm and environment transform tools to a supposed "AGI", sure does seem like cheating to me. Sad part is, it's a cheat that only works on environments where pathfinding is a major part. And when it doesn't have those clues it bombs on everything. I guess you really want to love the current SOTA LLMs. It's a shame they're dumb af. Have a great day.

>Adding a path finding algorithm and environment transform tools to a supposed "AGI", sure does seem like cheating to me. You would need all that if you, a human wanted any chance of solving this benchmark in the format LLMs are given. The funny thing about this benchmark is that we don't even know how solvable it is, because the baseline is tested with radically different inputs. >I guess you really want to love the…

Really tired of you making up stuff about this. The baseline and entire benchmark evaluation is clearly defined, with a statistically sound number of participants for the baseline using the same consistent deterministic environments to perform evaluation. The fact you don't like where the "human performance" line was drawn or how the scale is derived is not the same as the benchmark being tested with "radically different inputs". Clearly you would rather hype AI than critically advance it. So I won't waste time with someone who is clearly not posting in good faith.

Byebye now.

Re: ARC-AGI-3

#379

Earlier quoted context omitted.

The purpose is to benchmark both generality and intelligence. "Making up for" a poor score on one test with an excellent score on another would be the opposite of generality. There's a ceiling based on how consistent the performance is across all tasks.

>"Making up for" a poor score on one test with an excellent score on another would be the opposite of generality. Really ? This happens plenty with human testing. Humans aren't general ? The score is convoluted and messy. If the same score can say materially different things about capability then that's a bad scoring methodology. I can't believe I have to spell this out but it seems critical thinking goes out the win…

Just because humans are usually tested in a particular way that allows them to make up for a lack of generality with an outstanding performance in their specialization doesn't mean that is a good way to test generalization itself.

Apparently someone here doesn't know how outliers affect a mean. Or, for that matter, have any clue about the purpose of the ARC-AGI benchmark.

For anyone who is interested in critical thinking, this paper describes the original motivation behind the ARC benchmarks:

https://arxiv.org/abs/1911.01547

Re: ARC-AGI-3

#380

Earlier quoted context omitted.

>Adding a path finding algorithm and environment transform tools to a supposed "AGI", sure does seem like cheating to me. You would need all that if you, a human wanted any chance of solving this benchmark in the format LLMs are given. The funny thing about this benchmark is that we don't even know how solvable it is, because the baseline is tested with radically different inputs. >I guess you really want to love the…

Really tired of you making up stuff about this. The baseline and entire benchmark evaluation is clearly defined, with a statistically sound number of participants for the baseline using the same consistent deterministic environments to perform evaluation. The fact you don't like where the "human performance" line was drawn or how the scale is derived is not the same as the benchmark being tested with "radically diffe…

Humans and LLMs are not seeing the benchmark in the same format. What's made up about that ? Can you solve this in the JSON format ?

Look man, don't reply if you don't want to.

Post reply on HN