Live data from Hacker News

OpenAI's GPT-6 Astra on ARC-AGI-3

arcprize.org

31–40 of 153 posts

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#31
post #17

Is solving a snake like puzzle game in the least number of moves really what defines intelligence?

No, but figuring out that you're playing a snake-like puzzle game at all in an extremely general input domain and then solving it in the least number of moves definitely feels like evidence of intelligence.

You forget the benchmark. The human subjects were told they were being timed. If you believe the lowest time is the primary metric you will absolutely trial and error at speed instead of meticulously plan out your moves to minimize that metric.

LLMs are not timed and given that it costs tens of thousands of dollars to run this test they're not optimizing for speed.

So you've got a deceptive test, with one metric being told to humans and not applied to LLM and a hidden metric humans aren't aware of but LLMs are as the test.

This is flawed from the get go. It almost seems like this was deliberately setup to be able to claim AGI and superiority of LLMs

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#32
post #17

Is solving a snake like puzzle game in the least number of moves really what defines intelligence?

I'm also unclear as to how basic inferential logic puzzles spells out intelligence

I think if you summed up measures of intelligence as 'can it do basic symbolic logic in a chain with memory' then yes, you've now achieved the intelligence of an e. coli colony [0], congratulations

[0] https://journals.aps.org/prx/abstract/10.1103/PhysRevX.10.03...

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#33
post #17

Is solving a snake like puzzle game in the least number of moves really what defines intelligence?

You should read more on the ARC prize, it actually has a pretty long history. We're on the 3rd iteration because they keep getting saturated. If you look at the score history over time on ARC AGI 1, 2 and 3 it's pretty impressive.

https://arcprize.org/

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#35
post #17

Is solving a snake like puzzle game in the least number of moves really what defines intelligence?

There's currently a big market for figuring out ways to measure intelligence. With a particular interest in ways that humans can score much higher than LLMs. If you have some ideas please do share!

Why? Seems like benchmarks that closely mirror the tasks you'd want an LLM to help with would be a lot more useful than some general intelligence benchmark.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#36
post #17

Is solving a snake like puzzle game in the least number of moves really what defines intelligence?

1.5 years ago Gemini Pro 2.5 needed 1 page of thinking for every move in tic-tac-toe.

Playing tic-tac-toe or snake does not imply AGI, but is required to claim AGI.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#38

99.9% with the right harness? Ok, we're at AGI then. Prediction: We will now see the goalposts moved towards "well, a human costs less / is more efficient" - that will prevail for a few months until they come up with some other test that humans can do easily but is hard for the bots. This cycle will continue for ever and in 25 years, despite having hyper intelligent embodied robots or whatever, we'll still be arguing…

Then why is unemployment around 4%? You believe we have AGI and yet it can’t do anyone’s job?

Didn’t I just see a thing about how actual unemployment is at like 24% a few days ago?

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#39
post #17

Is solving a snake like puzzle game in the least number of moves really what defines intelligence?

It's pretty close to how we measure IQ. The standard test is basically a series of spatial puzzles.

I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the proxy for AGI, but then early LLMs could easily pass for a human in a casual conversation while clearly not matching human performance on most other tasks.

Since then, every benchmark we come up with, it turns out that an LLM can be fine-tuned to solve it while still clearly lacking something. They make very non-human mistakes, are easily tricked because they have a pretty tenuous grasp of reality, etc. But I think this just shows that AGI is a meaningless marketing term. We could as well be arguing if they have souls.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#40
post #23

Earlier quoted context omitted.

There's currently a big market for figuring out ways to measure intelligence. With a particular interest in ways that humans can score much higher than LLMs. If you have some ideas please do share!

I am no where close to qualified to do that. Hell, experts can't even define what intelligence is, much less define a test for it

It was defined in Animal Intelligence by George John Ramones in 1882 as "intelligence is the capacity to do the right thing at the right time. It is the ability to respond to the opportunities and challenges presented by a context"
Post reply on HN