Live data from Hacker News

OpenAI's GPT-6 Astra on ARC-AGI-3

arcprize.org

21–30 of 132 posts

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#21
post #17

Is solving a snake like puzzle game in the least number of moves really what defines intelligence?

There's currently a big market for figuring out ways to measure intelligence. With a particular interest in ways that humans can score much higher than LLMs. If you have some ideas please do share!

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#23
post #17

Is solving a snake like puzzle game in the least number of moves really what defines intelligence?

There's currently a big market for figuring out ways to measure intelligence. With a particular interest in ways that humans can score much higher than LLMs. If you have some ideas please do share!

I am no where close to qualified to do that. Hell, experts can't even define what intelligence is, much less define a test for it

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#25
post #23

Earlier quoted context omitted.

There's currently a big market for figuring out ways to measure intelligence. With a particular interest in ways that humans can score much higher than LLMs. If you have some ideas please do share!

I am no where close to qualified to do that. Hell, experts can't even define what intelligence is, much less define a test for it

This is like arguing about whether a hot dog is a sandwich (of course it is) or whether the chicken or the egg was first (obviously the egg since all chickens come from eggs). Intelligence is just problem solving in the context of self-awareness. Machines don't have it and never will but they can simulate the process given inputs. You can argue whether humans and animals truly possess self-awareness and in what degree, but the definition of intelligence is as simple as the hot dog debate.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#26

99.9% with the right harness? Ok, we're at AGI then. Prediction: We will now see the goalposts moved towards "well, a human costs less / is more efficient" - that will prevail for a few months until they come up with some other test that humans can do easily but is hard for the bots. This cycle will continue for ever and in 25 years, despite having hyper intelligent embodied robots or whatever, we'll still be arguing…

Then why is unemployment around 4%? You believe we have AGI and yet it can’t do anyone’s job?

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#27
post #17

Is solving a snake like puzzle game in the least number of moves really what defines intelligence?

No, but figuring out that you're playing a snake-like puzzle game at all in an extremely general input domain and then solving it in the least number of moves definitely feels like evidence of intelligence.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#28
post #22
post #20

What are these numbers? Why do they add up to a few hundred thousand dollars? Who paid for that? With what?

OpenAI provides API key with ~unlimited use?

So, you’re telling me I need to start a benchmark as a side gig to get a bunch of free compute.

Astra please create a benchmark that’s favorable to your reasoning skills with a human interface but don’t make the score too attainable add some small issues that keep you below 100% to look sensible and to keep my evaluation metric side gig going.

Alignment++

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#29
post #11

Earlier quoted context omitted.

DeepSeek v4 Flash recently had a similar "more reasoning is cheaper" curve. It's a fun counterintuition.

What is the intuition. Higher quality turns due to more reasoning results in significantly fewer turns taken?

Yes, in theory.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#30

what the hell is that score/cost curve lol

DeepSeek v4 Flash recently had a similar "more reasoning is cheaper" curve. It's a fun counterintuition.

It’s not that different than a lot of real world economies. Often paying for someone or something with better quality can reduce total costs. You have less failures, less mistakes, so on, so while the expertise or quality of the product is higher than cheaper solutions, they can be more reliable and over time ultimately cheaper.

The question I have is how far back that curve can go without relying on economies of scale to just drag all the points back to the left. And without overfitting a specific metric that I don’t need (like this test).

Post reply on HN