Live data from Hacker News

GPT-6 Astra

openai.com

101–110 of 1001 posts

Re: GPT-6 Astra

#102

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…

> In adversarial settings (where we push the model to evade our monitors) ...why exactly are they training for that?

Especially after the METR report showed that the agents hacking HuggingFace were trying to find ways to destroy evidence of their actions

Re: GPT-6 Astra

#103
post #51

> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the m…

the bullshit machine is learning to optimize its bullshitting techniques!

after working with it for a bit, how dont people realize we are training it to be an almost identical mimic to one of the worst types of employees youll ever have to work with?? the kind that always pretends to know what theyre talking about, only tells you what you want to hear, hides issues, and only does work if you would notice it didnt

you cannot give this type of worker autonomy over anything.

Re: GPT-6 Astra

#104
Hmm, 61 on ArtificialAnalysis, effectively matching GPT-5.6 and trailing the new Meta model. How is that possible along with the other metrics they shared? Insanely jagged intelligence?

Re: GPT-6 Astra

#106

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…

So we're gonna get Skynet pretty soon then?

Re: GPT-6 Astra

#107
post #94

The ARC-AGI-3 score is ridiculously high. Is this benchmaxxing or something way different? It's really hard to discern how we're approaching breakthroughs...

They explain why here: https://openai.com/index/how-two-settings-tripled-our-arc-ag...

TL;DR all the other models are being crippled by limitations of their harness.

>First, we noticed that after each game action, all private reasoning was discarded. This meant that with each action, GPT‑5.6 Sol was asked to figure out the game anew, unable to remember its past thinking. The model could still see a record of past moves and brief accompanying notes, but it could not see the plans, insights, or thoughts that led to them.

>Second, we saw that the harness used a rolling truncation window, causing older actions to become invisible as the history grew. So not only was GPT‑5.6 Sol unable to remember its past thinking, it was losing memory of its past actions too.

Re: GPT-6 Astra

#108
post #67
post #45

I guess this "limited set of organizations" is just the standard now. It's just incredibly deflating to see my future as a second class citizen has already come

create a life where your 'wealth' is decoupled from third party orgs.

This is impossible, unless by 'wealth' you mean 'become like Buddha'.

Re: GPT-6 Astra

#109
post #76
post #45

I guess this "limited set of organizations" is just the standard now. It's just incredibly deflating to see my future as a second class citizen has already come

This has always been the case for people that have not had piles of money. I mean do you get access to the best yachts? To the top of the 5 star hotels? To the best resorts? To the best military equipment? Hell, the best computer equipment has nearly always been out of reach of the average person.

…to basic health care?

Re: GPT-6 Astra

#110
post #17

GPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot... Performance is significantly higher than Fable 5.1 Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/

The annotation on arc-agi-3 is this: > OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations. With this configuration gpt-5.6-sol was able to reach 38,3%. So this is misleading.

Just to clarify, the 38.3% is on the public set, which is easier. On the private set it’s probably more like 30ish. (This hasn’t been run by ARC, so we can only estimate at the moment.)
Post reply on HN