We have such great AI and cannot keep a static site up?
GPT-6 Astra
91–100 of 1001 posts
Re: GPT-6 Astra
#92Re: GPT-6 Astra
#93GPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot... Performance is significantly higher than Fable 5.1 Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/
Is the ARC-AGI-3 score with their custom harness? I'm guessing that is what the footnote is for? (per https://openai.com/index/how-two-settings-tripled-our-arc-ag... )
ARC is reporting our score on their official leaderboard here: https://arcprize.org/leaderboard
A fair ding is that the comparison with Sol is not apples-to-apples (which we footnoted in the blog), but it's because we don’t have that data. I expect Sol would score roughly 30% with the responses API harness, so the Astra improvement is more like 30% -> 99% than 8% -> 99%. Still pretty good!
(I coauthored the linked blog post)
Re: GPT-6 Astra
#94> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…
> In adversarial settings (where we push the model to evade our monitors) ...why exactly are they training for that?
Re: GPT-6 Astra
#95You should know: AA index is only 61. Pretty surprised it’s that low.
And in the past, gemini 3 pro was rated as high as opus 4.5 and the like
Their AA Intelligence Index is just simply not indicative of whatever I care about, that's for sure.
Re: GPT-6 Astra
#96> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…
Sounds fun. As fun as their press release claiming it is the most safety aligned model ever.
Maybe they don't know themselves what's really going on. We are all in the interesting times gang now.
Re: GPT-6 Astra
#97We have such great AI and cannot keep a static site up?
Sometimes being the busiest site in the world for a few moments is difficult.
Re: GPT-6 Astra
#98Not on Azure? If so, that's a big deal.
Re: GPT-6 Astra
#99ARC AGI-3 saturated by Astra! https://arcprize.org/leaderboard
Re: GPT-6 Astra
#100> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…
> In adversarial settings (where we push the model to evade our monitors) ...why exactly are they training for that?