Live data from Hacker News

OpenAI's GPT-6 Astra on ARC-AGI-3

arcprize.org

141–150 of 150 posts

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#141
post #89

Earlier quoted context omitted.

We have always been very aware that IQ does not measure all skills - nonetheless, it does measure one. Nobody says that the IQ test would "measure intelligence". We know it does test a form of it.

I said it's a bad measure of intelligence and it seems that we agree.

That would be odd framing: it's a good test for a form of intelligence and other forms of intelligence still require good tests. It remains a good metric - for its specific thing; those who believe it to be the whole metric are naïve. It is almost necessary though not really sufficient.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#142
post #99
post #64

Earlier quoted context omitted.

> basic inferential logic puzzles spells out It spells out a form of intelligence - some can and some cannot. Those puzzles are an abstraction of a skill which is thought to be exportable in other domains.

so give it an IQ test and call it AGI. these weird little puzzles are grounded in no research with no replication or mechanistic chain to practical use

> so give it an IQ test and call it

more intelligent (in some specific dimensions of Intelligence) than what gets worse marks. Yes, pretty informally - but still notably.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#144

Anything you can verify to be right or wrong can be done by a model. All benchmarks will be saturated

You mean any repeatable benchmark will be saturated. The problem is that there is a huge perverse incentive. The intelligence is in the training layer not in the model parameters, but the intelligence is really good at remembering things, so if you let it take the test, it can RL it.

this is a short term problem, over 20 years benchmark gaming will be a blip

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#146
post #72

“AGI” never made sense to me. It’s a purely marketing term right? I’ve ignored it thinking it would go away, but it keeps coming up. I get that consciousness differs from intelligence and that our waking awareness of life is a complete mystery. Knowledge and thus intelligence however I consider as actively being solved by these large ML models. That is, with the right combination of machinery and know-how, you’ll get…

AGI is the term invented because arguments about what AI meant had gotten annoying. Originally there was no distinction and people thought "AI" would mean human level intelligence. Chess and conversations and robotics and math all in one package. Then games and classification and some robotics got solved and called AI, but that didn't solve math or conversation or online learning or a host of other things, so AGI was coined to refer to most of the whole package, virtually all the capabilities you'd need to replace humans intellectually. Now we're quibbling about whether AGI includes robotics or full real-world physical agents or something less.

The consensus now seems to be that once you've got human-level intelligence and planning and executive function then you get recursive self-improvement that can eventually autonomously solve the robotics and world-modeling and other portions of human-equivalence.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#147
post #72

“AGI” never made sense to me. It’s a purely marketing term right? I’ve ignored it thinking it would go away, but it keeps coming up. I get that consciousness differs from intelligence and that our waking awareness of life is a complete mystery. Knowledge and thus intelligence however I consider as actively being solved by these large ML models. That is, with the right combination of machinery and know-how, you’ll get…

[dead]

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#148

Earlier quoted context omitted.

The average human would have average intelligence

What's the metric on "average human" from a global population of > 8 billion people .. and why would they score mid on a standardised IQ test skewed toward western education / culture?

Who said anything about IQ tests?

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#149
post #135
post #54

Earlier quoted context omitted.

Pure software benchmarks might be getting saturated, but physical ones aren't. Let LLM control a physical robot to perform tasks that average human can do.

Not an llm, but: https://youtu.be/SzvvhRPj6eU?is=4ztgwim0GmjE1h-Y That (or a near future one) combined with an llm would be insane

That is a marketing teleoperation video, heavy with cuts.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#150
post #74

Earlier quoted context omitted.

> No, it is just because they have difficulties at the bench. I'll put it in another way. A "gifted kid" can be measured incredibly well on an IQ test, but fail miserably at incredibly normal but very difficult tasks such as consoling someone for their loss and managing family crisis. This is a clear example where an IQ measure doesn't translate to a person being capable of meaningfully changing their environments fo…

I don’t follow your logic. “The gifted person is highly intelligent, just not good at deadlifting 500kg” does not diminish the 500kg deadlift and I wouldn’t expect an IQ test to measure it.

Deadlifting 500kg is not a form of intelligence (or it could be under certain scenarios). I carefully chose specific tasks for my comment, because those tasks do reflect a kind of intelligence that's not measured by IQ.

I think the key here is to be wary of measurements that promise to capture the whole of what we consider intelligence (ie what people think of with IQs).

Post reply on HN