Is solving a snake like puzzle game in the least number of moves really what defines intelligence?
It's pretty close to how we measure IQ. The standard test is basically a series of spatial puzzles. I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the pro…
OpenAI's GPT-6 Astra on ARC-AGI-3
51–60 of 127 posts
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#5299.9% with the right harness? Ok, we're at AGI then. Prediction: We will now see the goalposts moved towards "well, a human costs less / is more efficient" - that will prevail for a few months until they come up with some other test that humans can do easily but is hard for the bots. This cycle will continue for ever and in 25 years, despite having hyper intelligent embodied robots or whatever, we'll still be arguing…
We were already there with Opus 5 -> https://www.linkedin.com/pulse/nvidias-harness-hit-100-arc-a...
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#53Is solving a snake like puzzle game in the least number of moves really what defines intelligence?
There's currently a big market for figuring out ways to measure intelligence. With a particular interest in ways that humans can score much higher than LLMs. If you have some ideas please do share!
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#54Is solving a snake like puzzle game in the least number of moves really what defines intelligence?
It's pretty close to how we measure IQ. The standard test is basically a series of spatial puzzles. I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the pro…
Let LLM control a physical robot to perform tasks that average human can do.
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#55Earlier quoted context omitted.
It's pretty close to how we measure IQ. The standard test is basically a series of spatial puzzles. I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the pro…
I don't think IQ is a good measure for intelligence at all. Neither dolphins or octopuses can solve IQ tests.
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#56Earlier quoted context omitted.
It's pretty close to how we measure IQ. The standard test is basically a series of spatial puzzles. I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the pro…
When I was 18, my high school girlfriend took me to the local Mensa chapter’s New Year’s party because her mother was a member and she was used to hanging out there. It was a useful lesson that whatever IQ tests measure, it is completely devoid of value or interest to me.
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#57Is solving a snake like puzzle game in the least number of moves really what defines intelligence?
It's pretty close to how we measure IQ. The standard test is basically a series of spatial puzzles. I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the pro…
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#58Anything you can verify to be right or wrong can be done by a model. All benchmarks will be saturated
Disagree. Examples: - predict a coinflip: easy to verify, hard to learn - earn $100: easy to verify, hard to learn - increase paid subscriptions in an A/B test: easy to verify, hard to learn I won't get into it, but there are many properties beyond verifiability that are needed to saturate a benchmark.
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#59Earlier quoted context omitted.
these just need more compute: - earn $100: easy to verify, hard to learn - increase paid subscriptions in an A/B test: easy to verify, hard to learn but we both know these examples go against the spirit of my point
Perhaps, but I think a bigger problem than lack of compute is the cost of rewards. Games like Chess and Go were solved long before self-driving, partly because it's incredibly cheap to acquire the reward of a bad board game decision, relatively to how expensive it is to acquire the cost of a bad driving decision. With driving, acquiring the reward can cost you $20/hr for human supervisors to generate disengagements,…
also, you are underestimating how short a 10 year time frame is. we are close to self driving, the first neural net image model was in 2013. 13 years is a blink of an eye
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#60Earlier quoted context omitted.
I don't think IQ is a good measure for intelligence at all. Neither dolphins or octopuses can solve IQ tests.
Their input and output interfaces are too different from human's and they're not nearly as smart to take our IQ tests, but both dolphins and octopuses can solve complex puzzles tailored for their environment. Those puzzles are the whole reason scientists know that dolphins and octopuses are more intelligent than other animals.