I guess this "limited set of organizations" is just the standard now. It's just incredibly deflating to see my future as a second class citizen has already come
GPT-6 Astra
81–90 of 1001 posts
Re: GPT-6 Astra
#82Re: GPT-6 Astra
#83I'm seeing reporting it gets 98.6% on ARC-AGI3[1] (previously like 30% with Fable) https://venturebeat.com/technology/welcome-to-the-agi-era-op...
This is with the caveat that OpenAI uses their own harness for this: > On ARC-AGI-3, GPT-6 Astra was run with our responses API harness , which changes two settings to better match real-world performance. The changes do not specifically target ARC-AGI-3.
Re: GPT-6 Astra
#84You should know: AA index is only 61. Pretty surprised it’s that low.
Re: GPT-6 Astra
#85The docs page has a bunch more interesting details, including for example async tool calling!
Re: GPT-6 Astra
#86$10 per million input tokens and $50 per million output tokens sol is $4 / $20
2.5x more expensive than Sol. Can expect 2.5x more usage in Codex subscription. Sol is already brutal (even after their recent fixes, it's just a token-hungry model: I go through a full 20x account per day, on Sol Med/High standard speed, with ~2 threads). I hope the efficiency gains are true, since their token efficiency claims for Sol were bullshit.
Do you use the official harness? OpenAI's models are generally best in class for token efficiency. It seems to me like they push for that much more than their competitors.
Re: GPT-6 Astra
#87> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…
...why exactly are they training for that?
Re: GPT-6 Astra
#88We have such great AI and cannot keep a static site up?
Re: GPT-6 Astra
#89Re: GPT-6 Astra
#90GPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot... Performance is significantly higher than Fable 5.1 Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/
Is the ARC-AGI-3 score with their custom harness? I'm guessing that is what the footnote is for? (per https://openai.com/index/how-two-settings-tripled-our-arc-ag... )
ARC is reporting our score on their official leaderboard here: https://arcprize.org/leaderboard
A fair ding is that the comparison with Sol is not apples-to-apples (which we footnoted in the blog), but it's because we don’t have that data. I expect Sol would score roughly 30% with the responses API harness, so the Astra improvement is more like 30% -> 99% than 8% -> 99%. Still pretty good!
(I coauthored the linked blog post)