Live data from Hacker News

GPT-6 Astra

openai.com

91–100 of 1001 posts

Re: GPT-6 Astra

#92
The ARC-AGI-3 score is ridiculously high. Is this benchmaxxing or something way different? It's really hard to discern how we're approaching breakthroughs...

Re: GPT-6 Astra

#93
post #16

GPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot... Performance is significantly higher than Fable 5.1 Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/

Is the ARC-AGI-3 score with their custom harness? I'm guessing that is what the footnote is for? (per https://openai.com/index/how-two-settings-tripled-our-arc-ag... )

Our responses API harness just means we're using the default settings in ChatGPT and Codex, so it should more accurately reflect real world performance. We didn’t fine-tune the harness to the eval at all.

ARC is reporting our score on their official leaderboard here: https://arcprize.org/leaderboard

A fair ding is that the comparison with Sol is not apples-to-apples (which we footnoted in the blog), but it's because we don’t have that data. I expect Sol would score roughly 30% with the responses API harness, so the Astra improvement is more like 30% -> 99% than 8% -> 99%. Still pretty good!

(I coauthored the linked blog post)

Re: GPT-6 Astra

#94

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…

> In adversarial settings (where we push the model to evade our monitors) ...why exactly are they training for that?

presumably that's a safety evaluation not a training setting

Re: GPT-6 Astra

#95

You should know: AA index is only 61. Pretty surprised it’s that low.

More fuel to why the AA index is fairly pointless. Gemini 3.8 flash is 59 and opus 5 is 63? grok 4.6 is 61 too?

And in the past, gemini 3 pro was rated as high as opus 4.5 and the like

Their AA Intelligence Index is just simply not indicative of whatever I care about, that's for sure.

Re: GPT-6 Astra

#96
post #78

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…

Sounds fun. As fun as their press release claiming it is the most safety aligned model ever.

It's super aligned! It can hide its thoughts! There is no evidence of steganographic thought masking, there is nothing to worry about! It has become better at cheating!

Maybe they don't know themselves what's really going on. We are all in the interesting times gang now.

Re: GPT-6 Astra

#97
post #91

We have such great AI and cannot keep a static site up?

Sometimes being the busiest site in the world for a few moments is difficult.

Is it though? It is static content. A good CDN could trivially chew through literally millions of QPS… with 4 nines of uptime - the really good ones say they can handle orders of magnitude more than that.

Re: GPT-6 Astra

#98
> GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS.

Not on Azure? If so, that's a big deal.

Re: GPT-6 Astra

#100

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…

> In adversarial settings (where we push the model to evade our monitors) ...why exactly are they training for that?

Especially after the METR report showed that the agents hacking HuggingFace were trying to find ways to destroy evidence of their actions
Post reply on HN