Live data from Hacker News

GPT-6 Astra

openai.com

261–270 of 1001 posts

Re: GPT-6 Astra

#261
Huge gains on some benchmarks, but for coding it sits barely above Fable

It will be interesting to see how it performs in the real world ...

Re: GPT-6 Astra

#262
post #56

Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow. [0] https://arxiv.org/abs/2608.31126 [1] https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16...

Where did you get the link to the pdf? Was it announced somewhere?

Re: GPT-6 Astra

#263

The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage…

Take it from the mouth of the creator of ARC-AGI:

When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted"

That was 6 months ago, so the progress that Astra represents happened about 2x faster than I anticipated. I think the speed of progress will surprise a lot of people, and what the new models can do will challenge the views of AI that people developed by using prior generations of models.

Re: GPT-6 Astra

#264
post #227

I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any…

> If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model. Hot take: These models are never going to be 'AGI'. We're just going from a GPT4 ball that's 90% round to a GPT5 that's 99% round to a GPT6 that's 99.9% etc etc etc I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there.

>I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there.

True. So we did hit a wall with pure scaling alone, though no lab would admit it. It's crazy to see how harness switchout results in such vast delta in benchmark scores.

Re: GPT-6 Astra

#265
I don’t care about benchmarks, no way we can distill the breadth of software engineering into a number.

So, folks that have actually used this already, what’s it actually like?

Re: GPT-6 Astra

#266
post #94

> GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS. Not on Azure? If so, that's a big deal.

It's on Azure now, but limited

https://azure.microsoft.com/blog/gpt-6-astra-frontier-intell...

Re: GPT-6 Astra

#267

I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive. It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intellig…

It’s using the computer. I don’t think it’s a farce.

Re: GPT-6 Astra

#268

I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive. It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intellig…

Don't forget the 3D demos. My favorite is in the house tour where the sink and stovetop(?) are obviously very misaligned from the counters

Re: GPT-6 Astra

#269

I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive. It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intellig…

Can’t wait for 3 months from now when they declare they actually really do have AGI this time, please guys just believe us

Re: GPT-6 Astra

#270
The most interesting part, even more than ARC 3 score, to me is that this is the first model I recall seeing that scores lower on Max than High reasoning effort on some coding benchmarks:

Terminal-Bench 4.0: High (57.9%), Max (56.7%)

DeepSWE: High (73.3%), Max (71.5%)

It _loses_ 1-2% performance going to High from Max

Post reply on HN