Earlier quoted context omitted.
We have always been very aware that IQ does not measure all skills - nonetheless, it does measure one. Nobody says that the IQ test would "measure intelligence". We know it does test a form of it.
I said it's a bad measure of intelligence and it seems that we agree.
OpenAI's GPT-6 Astra on ARC-AGI-3
141–150 of 152 posts
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#142Earlier quoted context omitted.
> basic inferential logic puzzles spells out It spells out a form of intelligence - some can and some cannot. Those puzzles are an abstraction of a skill which is thought to be exportable in other domains.
so give it an IQ test and call it AGI. these weird little puzzles are grounded in no research with no replication or mechanistic chain to practical use
more intelligent (in some specific dimensions of Intelligence) than what gets worse marks. Yes, pretty informally - but still notably.
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#143Are we sure an Astra hacker swarm didn't compromise arcprize.org's servers and exfiltrate the private eval set in order to achieve that 99%?
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#144Anything you can verify to be right or wrong can be done by a model. All benchmarks will be saturated
You mean any repeatable benchmark will be saturated. The problem is that there is a huge perverse incentive. The intelligence is in the training layer not in the model parameters, but the intelligence is really good at remembering things, so if you let it take the test, it can RL it.
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#145Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#146“AGI” never made sense to me. It’s a purely marketing term right? I’ve ignored it thinking it would go away, but it keeps coming up. I get that consciousness differs from intelligence and that our waking awareness of life is a complete mystery. Knowledge and thus intelligence however I consider as actively being solved by these large ML models. That is, with the right combination of machinery and know-how, you’ll get…
The consensus now seems to be that once you've got human-level intelligence and planning and executive function then you get recursive self-improvement that can eventually autonomously solve the robotics and world-modeling and other portions of human-equivalence.
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#147“AGI” never made sense to me. It’s a purely marketing term right? I’ve ignored it thinking it would go away, but it keeps coming up. I get that consciousness differs from intelligence and that our waking awareness of life is a complete mystery. Knowledge and thus intelligence however I consider as actively being solved by these large ML models. That is, with the right combination of machinery and know-how, you’ll get…
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#148Earlier quoted context omitted.
The average human would have average intelligence
What's the metric on "average human" from a global population of > 8 billion people .. and why would they score mid on a standardised IQ test skewed toward western education / culture?
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#149Earlier quoted context omitted.
Pure software benchmarks might be getting saturated, but physical ones aren't. Let LLM control a physical robot to perform tasks that average human can do.
Not an llm, but: https://youtu.be/SzvvhRPj6eU?is=4ztgwim0GmjE1h-Y That (or a near future one) combined with an llm would be insane
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#150Earlier quoted context omitted.
> No, it is just because they have difficulties at the bench. I'll put it in another way. A "gifted kid" can be measured incredibly well on an IQ test, but fail miserably at incredibly normal but very difficult tasks such as consoling someone for their loss and managing family crisis. This is a clear example where an IQ measure doesn't translate to a person being capable of meaningfully changing their environments fo…
I don’t follow your logic. “The gifted person is highly intelligent, just not good at deadlifting 500kg” does not diminish the 500kg deadlift and I wouldn’t expect an IQ test to measure it.
I think the key here is to be wary of measurements that promise to capture the whole of what we consider intelligence (ie what people think of with IQs).