I guess this "limited set of organizations" is just the standard now. It's just incredibly deflating to see my future as a second class citizen has already come
GPT-6 Astra
81–90 of 1001 posts
Re: GPT-6 Astra
#82> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the m…
Re: GPT-6 Astra
#83I guess this "limited set of organizations" is just the standard now. It's just incredibly deflating to see my future as a second class citizen has already come
Re: GPT-6 Astra
#84Re: GPT-6 Astra
#85I'm seeing reporting it gets 98.6% on ARC-AGI3[1] (previously like 30% with Fable) https://venturebeat.com/technology/welcome-to-the-agi-era-op...
This is with the caveat that OpenAI uses their own harness for this: > On ARC-AGI-3, GPT-6 Astra was run with our responses API harness , which changes two settings to better match real-world performance. The changes do not specifically target ARC-AGI-3.
Re: GPT-6 Astra
#86You should know: AA index is only 61. Pretty surprised it’s that low.
Re: GPT-6 Astra
#87The docs page has a bunch more interesting details, including for example async tool calling!
Re: GPT-6 Astra
#88I guess this "limited set of organizations" is just the standard now. It's just incredibly deflating to see my future as a second class citizen has already come
Oh please. They do closed betas - hardly makes you a "second class citizen".
Re: GPT-6 Astra
#89$10 per million input tokens and $50 per million output tokens sol is $4 / $20
2.5x more expensive than Sol. Can expect 2.5x more usage in Codex subscription. Sol is already brutal (even after their recent fixes, it's just a token-hungry model: I go through a full 20x account per day, on Sol Med/High standard speed, with ~2 threads). I hope the efficiency gains are true, since their token efficiency claims for Sol were bullshit.
Do you use the official harness? OpenAI's models are generally best in class for token efficiency. It seems to me like they push for that much more than their competitors.
Re: GPT-6 Astra
#90> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…
...why exactly are they training for that?