Live data from Hacker News

GPT-6 Astra

openai.com

291–300 of 1001 posts

Re: GPT-6 Astra

#291
post #247

I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any…

> If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model. Hot take: These models are never going to be 'AGI'. We're just going from a GPT4 ball that's 90% round to a GPT5 that's 99% round to a GPT6 that's 99.9% etc etc etc I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there.

>I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there.

True. So we did hit a wall with pure scaling alone, though no lab would admit it. It's crazy to see how harness switchout results in such vast delta in benchmark scores.

Re: GPT-6 Astra

#292
I don’t care about benchmarks, no way we can distill the breadth of software engineering into a number.

So, folks that have actually used this already, what’s it actually like?

Re: GPT-6 Astra

#293
post #98

> GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS. Not on Azure? If so, that's a big deal.

It's on Azure now, but limited

https://azure.microsoft.com/blog/gpt-6-astra-frontier-intell...

Re: GPT-6 Astra

#294

I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive. It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intellig…

It’s using the computer. I don’t think it’s a farce.

Re: GPT-6 Astra

#295

I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive. It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intellig…

Don't forget the 3D demos. My favorite is in the house tour where the sink and stovetop(?) are obviously very misaligned from the counters

Re: GPT-6 Astra

#296

I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive. It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intellig…

Can’t wait for 3 months from now when they declare they actually really do have AGI this time, please guys just believe us

Re: GPT-6 Astra

#297
The most interesting part, even more than ARC 3 score, to me is that this is the first model I recall seeing that scores lower on Max than High reasoning effort on some coding benchmarks:

Terminal-Bench 4.0: High (57.9%), Max (56.7%)

DeepSWE: High (73.3%), Max (71.5%)

It _loses_ 1-2% performance going to High from Max

Re: GPT-6 Astra

#298
post #247

I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any…

> If this is truly AGI (subject to one's definition of AGI still) Scoring well in a benchmark that's called AGI does not make an LLM AGI.

But they declared it...

Re: GPT-6 Astra

#299

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

If your definition of AGI involves copy/pasting code and doing well in some made up benchmark, then probably AGI is close

Re: GPT-6 Astra

#300

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

It has to pass the Turing test

LLMs started meaningfully passing the Turing test a year or two ago, around GPT-4.5. Is there another version or bar for "passing" you're looking for?

[0] https://arxiv.org/pdf/2503.23674

Post reply on HN