GPT-6 Astra
221–230 of 1001 posts
Re: GPT-6 Astra
#222Re: GPT-6 Astra
#223> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…
I really wish it was called chain of instruction. Because it's definitely not thought.
Re: GPT-6 Astra
#224ARC AGI-3 saturated by Astra! https://arcprize.org/leaderboard
But scored less on V2 and V1 ... too much overfitting?
Re: GPT-6 Astra
#225> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…
I really wish it was called chain of instruction. Because it's definitely not thought.
Re: GPT-6 Astra
#226Oh brotha, here we go again, it's so over again, as every week nowadays
But thank you for spending other peoples money to give us the tech regardless!
Re: GPT-6 Astra
#227Re: GPT-6 Astra
#228I guess this "limited set of organizations" is just the standard now. It's just incredibly deflating to see my future as a second class citizen has already come
Re: GPT-6 Astra
#229Finally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training. I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing. Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally…
yeah i'm wondering the same way... especially in light of the 20x debacle (where we found that 20x of Max vs 5x only applies to the 5hr limit, not the weekly limit, whereas OpenAI's 20x actually is 20x overall). Also Opus 5 has been really tough to work with. I can't understand half of what it says, it's just so damn obscure.
Re: GPT-6 Astra
#230Because in another dead language of antiquity, Sanskrit, it means "weapon". Which would be a bit too on-the-nose.