Live data from Hacker News

GPT-6 Astra

openai.com

81–90 of 1001 posts

Re: GPT-6 Astra

#81
post #45

I guess this "limited set of organizations" is just the standard now. It's just incredibly deflating to see my future as a second class citizen has already come

When Open AI announced that Astra was the first to reach the "Critical" level in cybersecurity it also said that advanced cyber capabilities are initially provided to a narrow circle of alpha testers like the US government and trusted organizations that Open AI doesn't name. To my mind the "Critical" level itself is an internal scale of Open AI its own Preparedness Framework and not an external audit.

Re: GPT-6 Astra

#83
post #34

I'm seeing reporting it gets 98.6% on ARC-AGI3[1] (previously like 30% with Fable) https://venturebeat.com/technology/welcome-to-the-agi-era-op...

This is with the caveat that OpenAI uses their own harness for this: > On ARC-AGI-3, GPT-6 Astra was run with our responses API harness , which changes two settings to better match real-world performance. The changes do not specifically target ARC-AGI-3.

This should be normalised and expected - the responses API harness allows it to use the custom compaction that is not allowed otherwise. It is entirely fair to allow OpenAI to use their own compaction algorithm..

Re: GPT-6 Astra

#84

You should know: AA index is only 61. Pretty surprised it’s that low.

I have some doubts about AA-index. For example Opus 5 (High) is at the same index value as Fable 5 (Max), that doesn't seem right.

Re: GPT-6 Astra

#86
post #10

$10 per million input tokens and $50 per million output tokens sol is $4 / $20

2.5x more expensive than Sol. Can expect 2.5x more usage in Codex subscription. Sol is already brutal (even after their recent fixes, it's just a token-hungry model: I go through a full 20x account per day, on Sol Med/High standard speed, with ~2 threads). I hope the efficiency gains are true, since their token efficiency claims for Sol were bullshit.

How do you manage to run out of tokens so quickly? I probably run more threads every working day, usually on medium, and I'm still below the 5x limits.

Do you use the official harness? OpenAI's models are generally best in class for token efficiency. It seems to me like they push for that much more than their competitors.

Re: GPT-6 Astra

#87

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…

> In adversarial settings (where we push the model to evade our monitors)

...why exactly are they training for that?

Re: GPT-6 Astra

#89
The ARC-AGI-3 score is ridiculously high. Is this benchmaxxing or something way different? It's really hard to discern how we're approaching breakthroughs...

Re: GPT-6 Astra

#90
post #16

GPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot... Performance is significantly higher than Fable 5.1 Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/

Is the ARC-AGI-3 score with their custom harness? I'm guessing that is what the footnote is for? (per https://openai.com/index/how-two-settings-tripled-our-arc-ag... )

Our responses API harness just means we're using the default settings in ChatGPT and Codex, so it should more accurately reflect real world performance. We didn’t fine-tune the harness to the eval at all.

ARC is reporting our score on their official leaderboard here: https://arcprize.org/leaderboard

A fair ding is that the comparison with Sol is not apples-to-apples (which we footnoted in the blog), but it's because we don’t have that data. I expect Sol would score roughly 30% with the responses API harness, so the Astra improvement is more like 30% -> 99% than 8% -> 99%. Still pretty good!

(I coauthored the linked blog post)

Post reply on HN