Live data from Hacker News

OpenAI's GPT-6 Astra on ARC-AGI-3

arcprize.org

41–50 of 137 posts

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#41

Anything you can verify to be right or wrong can be done by a model. All benchmarks will be saturated

Disagree.

Examples:

- predict a coinflip: easy to verify, hard to learn

- earn $100: easy to verify, hard to learn

- increase paid subscriptions in an A/B test: easy to verify, hard to learn

I won't get into it, but there are many properties beyond verifiability that are needed to saturate a benchmark.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#42
post #17

Is solving a snake like puzzle game in the least number of moves really what defines intelligence?

It's pretty close to how we measure IQ. The standard test is basically a series of spatial puzzles. I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the pro…

When I was 18, my high school girlfriend took me to the local Mensa chapter’s New Year’s party because her mother was a member and she was used to hanging out there.

It was a useful lesson that whatever IQ tests measure, it is completely devoid of value or interest to me.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#44
post #11

Earlier quoted context omitted.

DeepSeek v4 Flash recently had a similar "more reasoning is cheaper" curve. It's a fun counterintuition.

What is the intuition. Higher quality turns due to more reasoning results in significantly fewer turns taken?

Yep. In particular, ARC-AGI-3 is a series of games where if you fail, you keep trying again (until eventually hitting a timeout). So the sooner you succeed, the sooner you stop spending tokens retrying. If it was a benchmark where everyone got one attempt with no retries, you wouldn't see it bend backward.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#45

Anything you can verify to be right or wrong can be done by a model. All benchmarks will be saturated

Disagree. Examples: - predict a coinflip: easy to verify, hard to learn - earn $100: easy to verify, hard to learn - increase paid subscriptions in an A/B test: easy to verify, hard to learn I won't get into it, but there are many properties beyond verifiability that are needed to saturate a benchmark.

these just need more compute:

- earn $100: easy to verify, hard to learn

- increase paid subscriptions in an A/B test: easy to verify, hard to learn

but we both know these examples go against the spirit of my point

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#46

Earlier quoted context omitted.

Then why is unemployment around 4%? You believe we have AGI and yet it can’t do anyone’s job?

Didn’t I just see a thing about how actual unemployment is at like 24% a few days ago?

According to that interpretation, ~24% is one of the lowest ever.

https://www.lisep.org/tru

(I have not gone down the rabbit hole to understand how they achieve that 24% number)

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#47

99.9% with the right harness? Ok, we're at AGI then. Prediction: We will now see the goalposts moved towards "well, a human costs less / is more efficient" - that will prevail for a few months until they come up with some other test that humans can do easily but is hard for the bots. This cycle will continue for ever and in 25 years, despite having hyper intelligent embodied robots or whatever, we'll still be arguing…

They explain it here: https://openai.com/index/how-two-settings-tripled-our-arc-ag...

TLDR: The official ARC harness throws away old context and reasoning. No real-world harness is this bad, the model has to re-learn the game repeatedly. OpenAI basically just added standard compaction. Their harness is still "general".

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#49
post #43

The instant/no reasoning performed extremely well none 35.2%, $49,791 96.7%, $23,457 35.2% on the standard harness, that's above Opus 5 on high.

Since low scored much lower than none, and none scored ~ around medium, could none default to medium in the API? I don't think the new models can even have "instant" via API, unless they train them for that (there was one gpt5 variant called instant or something).

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#50

Earlier quoted context omitted.

Disagree. Examples: - predict a coinflip: easy to verify, hard to learn - earn $100: easy to verify, hard to learn - increase paid subscriptions in an A/B test: easy to verify, hard to learn I won't get into it, but there are many properties beyond verifiability that are needed to saturate a benchmark.

these just need more compute: - earn $100: easy to verify, hard to learn - increase paid subscriptions in an A/B test: easy to verify, hard to learn but we both know these examples go against the spirit of my point

Perhaps, but I think a bigger problem than lack of compute is the cost of rewards. Games like Chess and Go were solved long before self-driving, partly because it's incredibly cheap to acquire the reward of a bad board game decision, relatively to how expensive it is to acquire the cost of a bad driving decision. With driving, acquiring the reward can cost you $20/hr for human supervisors to generate disengagements, or $100k if you crash, or $30B if you crash the car into a person in a way that causes your company to collapse (e.g., Cruise).
Post reply on HN