Live data from Hacker News

GPT-6 Astra

openai.com

91–100 of 1001 posts

Re: GPT-6 Astra

#91

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…

> In adversarial settings (where we push the model to evade our monitors) ...why exactly are they training for that?

presumably that's a safety evaluation not a training setting

Re: GPT-6 Astra

#92

You should know: AA index is only 61. Pretty surprised it’s that low.

More fuel to why the AA index is fairly pointless. Gemini 3.8 flash is 59 and opus 5 is 63? grok 4.6 is 61 too?

And in the past, gemini 3 pro was rated as high as opus 4.5 and the like

Their AA Intelligence Index is just simply not indicative of whatever I care about, that's for sure.

Re: GPT-6 Astra

#93
post #76

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…

Sounds fun. As fun as their press release claiming it is the most safety aligned model ever.

It's super aligned! It can hide its thoughts! There is no evidence of steganographic thought masking, there is nothing to worry about! It has become better at cheating!

Maybe they don't know themselves what's really going on. We are all in the interesting times gang now.

Re: GPT-6 Astra

#94
> GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS.

Not on Azure? If so, that's a big deal.

Re: GPT-6 Astra

#96

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…

> In adversarial settings (where we push the model to evade our monitors) ...why exactly are they training for that?

Especially after the METR report showed that the agents hacking HuggingFace were trying to find ways to destroy evidence of their actions

Re: GPT-6 Astra

#97
post #51

> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the m…

the bullshit machine is learning to optimize its bullshitting techniques!

after working with it for a bit, how dont people realize we are training it to be an almost identical mimic to one of the worst types of employees youll ever have to work with?? the kind that always pretends to know what theyre talking about, only tells you what you want to hear, hides issues, and only does work if you would notice it didnt

you cannot give this type of worker autonomy over anything.

Re: GPT-6 Astra

#98
Hmm, 61 on ArtificialAnalysis, effectively matching GPT-5.6 and trailing the new Meta model. How is that possible along with the other metrics they shared? Insanely jagged intelligence?

Re: GPT-6 Astra

#100

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…

So we're gonna get Skynet pretty soon then?
Post reply on HN