Live data from Hacker News

OpenAI's GPT-6 Astra on ARC-AGI-3

arcprize.org

101–110 of 130 posts

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#101
post #54

Earlier quoted context omitted.

Pure software benchmarks might be getting saturated, but physical ones aren't. Let LLM control a physical robot to perform tasks that average human can do.

What makes you think this task won't fall quickly too?

It might, but if there's one thing that hasn't changed since 2022, it's that the models tend to ace the tasks where you have gobs of training data and where verification loops are fast and cheap... and they are not nearly as amazing elsewhere. If it's close enough, they can generalize, e.g. translate one programming language to another. But there's a pretty steep cliff past a certain distance.

Case in point: you had hundreds of millions of JPEGs to vacuum up and bitmap image generation is amazing. But if you ask them to recreate the same scene as vector art, they will struggle to generate a decent SVG. Like, kindergarten-style pelicans on bicycles are the state of the art. It should generalize seamlessly, but somehow, doesn't?

I think it will happen, just like self-driving cars are happening, but it will probably be a slow process.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#102

Earlier quoted context omitted.

Workforce participation is different than unemployment. Didn't click the link but I suspect that's the case with your 24%

24% sounds correct to me. Many of my friend who completed phd are currently unemployed .... And, some of them started doing random job like uber to make a living. If government don't step up and regulate outsourcing like 100% tax, big problems are coming up.

24.9% is damn close to the 24.8% average over the ast 25 years, for the thing it's measuring-- jobless, people working part-time or involuntarily, and workers earning less than $26,000. Since 2000, it's been over 30% for more years than it's been under 24%.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#103
post #88
post #80

Earlier quoted context omitted.

capability to match or surpass human intelligence across all conceivable cognitive tasks What human intelligence do you mean? Genius? Professional? Educated? Random person? “Dumb” person?

I think the distinction between these options probably doesn’t matter alll that much if the threshold you’re using is either “random person” or higher? Maybe bump it up to “random educated person”? They are still different concepts of course, but I imagine that once one is achieved, the others aren’t far off.

Ok, but a random educated person will not perform well on vast majority of specialized tasks where professionals operate. A model like Astra probably will beat random educated person performance on majority of specialized tasks. It’s getting close to the level of professionals in many domains, and to genius level on some (e.g. math).

I’m just trying to understand the implications of the current frontier model capabilities.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#104
post #72

“AGI” never made sense to me. It’s a purely marketing term right? I’ve ignored it thinking it would go away, but it keeps coming up. I get that consciousness differs from intelligence and that our waking awareness of life is a complete mystery. Knowledge and thus intelligence however I consider as actively being solved by these large ML models. That is, with the right combination of machinery and know-how, you’ll get…

The OpenAI charter defines it as: "highly autonomous systems that outperform humans at most economically valuable work" https://time.com/article/2026/08/26/openai-sam-altman-interv...

So vague it’s useless.

We cannot in a declarative sense define what is economically valuable work even now let alone into the future.

People take what they can get for pay. Very few individuals can demand a wage. The value of employment is obviously designed around that, not some arbitrary definition of “valuable”.

Of course an AI will accept $0/hr, it doesn’t mean it does the job.

Anyone who could accurately define the value of work would be wildly successful without having to try.

That is not a useful definition for me unfortunately.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#105
post #17

Is solving a snake like puzzle game in the least number of moves really what defines intelligence?

It's pretty close to how we measure IQ. The standard test is basically a series of spatial puzzles. I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the pro…

IQ tests are incredibly good at what they're designed for, which is discriminating relatively higher intelligence humans from lower intelligence humans. Also discriminating within a single human - they are routinely and reliably used to track cognitive decline.

For these purposes they are highly reliable (repeatable, internally consistent) and valid (correlate with ~everything to about the degree one would reasonably expect).

They were never designed for machines or non-human animals.

Nor were they designed for rare ranges of intelligence - these are by definition hard to create tests for, since it's hard to gather the sample sizes you need. So they work well for the middle ~98% of humans but can't discriminate well among the most profoundly intellectually disabled nor among true geniuses.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#106
post #93
post #72

“AGI” never made sense to me. It’s a purely marketing term right? I’ve ignored it thinking it would go away, but it keeps coming up. I get that consciousness differs from intelligence and that our waking awareness of life is a complete mystery. Knowledge and thus intelligence however I consider as actively being solved by these large ML models. That is, with the right combination of machinery and know-how, you’ll get…

> Knowledge and thus intelligence How can you conflate the two. > solving consciousness We are very much not interested in that. We just need a proper problem solver.

Unproven, but it could very well be some problems require consciousness to be solved.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#107
post #56

Earlier quoted context omitted.

What does having value or interest to you have to do with intelligence?

What does IQ have to do with intelligence? (Please forgive the flippant response. I believe it cuts to the core of what the parent was intending.)

Is it a serious question? IQ tries to quantify the positive correlation between the results of all intellectual tasks a person takes (AKA positive manifold).

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#108
post #89
post #74

Earlier quoted context omitted.

> No, it is just because they have difficulties at the bench. I'll put it in another way. A "gifted kid" can be measured incredibly well on an IQ test, but fail miserably at incredibly normal but very difficult tasks such as consoling someone for their loss and managing family crisis. This is a clear example where an IQ measure doesn't translate to a person being capable of meaningfully changing their environments fo…

We have always been very aware that IQ does not measure all skills - nonetheless, it does measure one. Nobody says that the IQ test would "measure intelligence". We know it does test a form of it.

I said it's a bad measure of intelligence and it seems that we agree.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#109
post #87

I like Erdos problems as a benchmark. Models continue to solve them, but a pretty tepid rate now that the low hanging fruit has been taken. From https://epoch.ai/latest/announcing-frontiermath-erdos > Only GPT-6 Astra solved anything: 2 of the 68 problems. It disproved problem 74 by finding a counterexample, at a cost of $218 and 15 hours of working time, and it proved problem 126, at a cost of $247 and 16 hours > Ac…

A very long tail of problems that weren't solved by humans? Sure.

It's a sarcastic take and I understand that you are probably talking about "spiky intelligence", but you've chosen unsolved problems as a measure of the progress yourself.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#110
post #104

Earlier quoted context omitted.

The OpenAI charter defines it as: "highly autonomous systems that outperform humans at most economically valuable work" https://time.com/article/2026/08/26/openai-sam-altman-interv...

So vague it’s useless. We cannot in a declarative sense define what is economically valuable work even now let alone into the future. People take what they can get for pay. Very few individuals can demand a wage. The value of employment is obviously designed around that, not some arbitrary definition of “valuable”. Of course an AI will accept $0/hr, it doesn’t mean it does the job. Anyone who could accurately define…

That's the challenge of defining intelligence. What is intelligence?

Do you want to take a stab at defining/quantifying it? I'm s afraid anything specific you can come up with will also be useless.

Post reply on HN