Live data from Hacker News

OpenAI's GPT-6 Astra on ARC-AGI-3

arcprize.org

51–60 of 148 posts

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#51
post #17

Is solving a snake like puzzle game in the least number of moves really what defines intelligence?

It's pretty close to how we measure IQ. The standard test is basically a series of spatial puzzles. I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the pro…

I don't think IQ is a good measure for intelligence at all. Neither dolphins or octopuses can solve IQ tests.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#52
post #8

99.9% with the right harness? Ok, we're at AGI then. Prediction: We will now see the goalposts moved towards "well, a human costs less / is more efficient" - that will prevail for a few months until they come up with some other test that humans can do easily but is hard for the bots. This cycle will continue for ever and in 25 years, despite having hyper intelligent embodied robots or whatever, we'll still be arguing…

We were already there with Opus 5 -> https://www.linkedin.com/pulse/nvidias-harness-hit-100-arc-a...

AFAICT nvda's result is on the 25 open problems, while this submission is on the "semi-private" set, ran by the arc people themselves.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#53
post #17

Is solving a snake like puzzle game in the least number of moves really what defines intelligence?

There's currently a big market for figuring out ways to measure intelligence. With a particular interest in ways that humans can score much higher than LLMs. If you have some ideas please do share!

Give away access to the model and go ask people from time to time if the model was of use to the person and if they were able to make the model work with them.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#54
post #17

Is solving a snake like puzzle game in the least number of moves really what defines intelligence?

It's pretty close to how we measure IQ. The standard test is basically a series of spatial puzzles. I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the pro…

Pure software benchmarks might be getting saturated, but physical ones aren't.

Let LLM control a physical robot to perform tasks that average human can do.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#55
post #51

Earlier quoted context omitted.

It's pretty close to how we measure IQ. The standard test is basically a series of spatial puzzles. I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the pro…

I don't think IQ is a good measure for intelligence at all. Neither dolphins or octopuses can solve IQ tests.

Their input and output interfaces are too different from human's and they're not nearly as smart to take our IQ tests, but both dolphins and octopuses can solve complex puzzles tailored for their environment. Those puzzles are the whole reason scientists know that dolphins and octopuses are more intelligent than other animals.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#56
post #42

Earlier quoted context omitted.

It's pretty close to how we measure IQ. The standard test is basically a series of spatial puzzles. I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the pro…

When I was 18, my high school girlfriend took me to the local Mensa chapter’s New Year’s party because her mother was a member and she was used to hanging out there. It was a useful lesson that whatever IQ tests measure, it is completely devoid of value or interest to me.

What does having value or interest to you have to do with intelligence?

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#57
post #17

Is solving a snake like puzzle game in the least number of moves really what defines intelligence?

It's pretty close to how we measure IQ. The standard test is basically a series of spatial puzzles. I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the pro…

Nit: Turing’s actual imitation game is a party game (like Werewolf/Mafia) and nobody’s even trying to win at that. The LLM’s will just tell you they’re an AI.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#58

Anything you can verify to be right or wrong can be done by a model. All benchmarks will be saturated

Disagree. Examples: - predict a coinflip: easy to verify, hard to learn - earn $100: easy to verify, hard to learn - increase paid subscriptions in an A/B test: easy to verify, hard to learn I won't get into it, but there are many properties beyond verifiability that are needed to saturate a benchmark.

Doesn't "saturated" mean that essentially there won't be any more progress in the benchmarch? Also of note is that two of your points only mean something on an occidental capitalist system.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#59

Earlier quoted context omitted.

these just need more compute: - earn $100: easy to verify, hard to learn - increase paid subscriptions in an A/B test: easy to verify, hard to learn but we both know these examples go against the spirit of my point

Perhaps, but I think a bigger problem than lack of compute is the cost of rewards. Games like Chess and Go were solved long before self-driving, partly because it's incredibly cheap to acquire the reward of a bad board game decision, relatively to how expensive it is to acquire the cost of a bad driving decision. With driving, acquiring the reward can cost you $20/hr for human supervisors to generate disengagements,…

yeah but I think you may be underestimating the amount of capital available for compute. if AGI is possible through some 5 trillion of expenditure on computers, there will be money for it.

also, you are underestimating how short a 10 year time frame is. we are close to self driving, the first neural net image model was in 2013. 13 years is a blink of an eye

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#60
post #55
post #51

Earlier quoted context omitted.

I don't think IQ is a good measure for intelligence at all. Neither dolphins or octopuses can solve IQ tests.

Their input and output interfaces are too different from human's and they're not nearly as smart to take our IQ tests, but both dolphins and octopuses can solve complex puzzles tailored for their environment. Those puzzles are the whole reason scientists know that dolphins and octopuses are more intelligent than other animals.

But we do know they're "intelligent" and also smart in an important capacity. So how gives we don't measure them by IQ? Because the IQ is not a good measure of intelligence or smarts.
Post reply on HN