Live data from Hacker News

GPT-6 Astra

openai.com

201–210 of 1001 posts

Re: GPT-6 Astra

#203

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…

I really wish it was called chain of instruction. Because it's definitely not thought.

Yeah, basically they are using more computation to explore the solution space before producing the final answer.

Re: GPT-6 Astra

#204
post #78

ARC AGI-3 saturated by Astra! https://arcprize.org/leaderboard

But scored less on V2 and V1 ... too much overfitting?

https://mvakde.github.io/blog/44-on-arc-1/ makes a good case that all the performance on the arc agi tests is overfitting, based on the fact that v1 performance did not translate directly to v2 performance

Re: GPT-6 Astra

#205

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…

I really wish it was called chain of instruction. Because it's definitely not thought.

"Chain of Thoughts" is a term from the title of a 2022 research paper "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models" (https://arxiv.org/abs/2201.11903), well before ChatGPT and the subsequent marketing hype. If anything, it's the most correct way to use the term.

Re: GPT-6 Astra

#206

Oh brotha, here we go again, it's so over again, as every week nowadays

I think Altman and amodei have a difficult time in understanding that you can have intelligent technology boxes but… it doesn’t change reality all that much.

But thank you for spending other peoples money to give us the tech regardless!

Re: GPT-6 Astra

#209
post #163

Finally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training. I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing. Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally…

yeah i'm wondering the same way... especially in light of the 20x debacle (where we found that 20x of Max vs 5x only applies to the 5hr limit, not the weekly limit, whereas OpenAI's 20x actually is 20x overall). Also Opus 5 has been really tough to work with. I can't understand half of what it says, it's just so damn obscure.

Could you share more about 5x/20x? I missed that

Re: GPT-6 Astra

#210
What does 'Astra' here mean? Surely they must be referring to the Latin word.

Because in another dead language of antiquity, Sanskrit, it means "weapon". Which would be a bit too on-the-nose.

Post reply on HN