Earlier quoted context omitted.
Do you have hard evidence of this assertion?
We can't have hard evidence. It's a SaaS and they own the code and the machine it runs on. So it may be a widespread hallucination. But there's no evidence of that either.
GPT-6 Astra, looped transformers, and hidden reasoning
111–120 of 151 posts
Re: GPT-6 Astra, looped transformers, and hidden reasoning
#112Earlier quoted context omitted.
NVFP4 would buy them a huge increase in capacity but I think it would be noticeable.
Why doesn't someone just try to measure this next time!?
Re: GPT-6 Astra, looped transformers, and hidden reasoning
#113Earlier quoted context omitted.
The most charitable explanation I can think of for this is something like regression to the mean. When a model is first released, there'll be a subset of users who, just by chance, sample the highest quality band of the distribution that answers their query. Some of them will rush over to social media and post about how amazing a model is. Over time, those users' mental model of responses will converge but they'll pe…
To add onto this, if you use a shiny new model and it gives you a turd, you're not going to tweet about it ("hey guys, look what I made with Astra! Nothing!"), and even if you do nobody is going to interact with it so it does poorly in the algorithm, because it has to compete with all the people using the new model to make something that looks impressive. Then people get tired of the magic trick and the logic flips.
Re: GPT-6 Astra, looped transformers, and hidden reasoning
#114Earlier quoted context omitted.
It's the same story every time OpenAI or Anthropic releases a new model. They are generous with compute for the first few days, and use maximum fidelity with uncompressed weights. Everything runs at its best to make a good first impression. But eventually they pare things back and the models perform a little worse.
The most charitable explanation I can think of for this is something like regression to the mean. When a model is first released, there'll be a subset of users who, just by chance, sample the highest quality band of the distribution that answers their query. Some of them will rush over to social media and post about how amazing a model is. Over time, those users' mental model of responses will converge but they'll pe…
Re: GPT-6 Astra, looped transformers, and hidden reasoning
#115Astra was insane until Monday but something happened on tuesday, now it feels like Sol. I grieve for the lost productivity but i hope they may give us the original Astra back.
It's the same story every time OpenAI or Anthropic releases a new model. They are generous with compute for the first few days, and use maximum fidelity with uncompressed weights. Everything runs at its best to make a good first impression. But eventually they pare things back and the models perform a little worse.
Re: GPT-6 Astra, looped transformers, and hidden reasoning
#116Earlier quoted context omitted.
Might be related to this announcement from Tibo on Sunday: > We've made some improvements that improve usage on the long tail for power users of Astra when logged in with your ChatGPT account. > No change in quality and a pure win that on the long tail can result in up to 3-4X less usage being drawn from the subscription. https://x.com/thsottiaux/status/2096717905614524491 ( https://xcancel.com/thsottiaux/status/2096…
It seems to me the people working at OAI may believe all other humans must be a little bit behind intellectually.
Point is, I really don't buy all the stories about a model suddenly being downgraded without at least a modicum of substance. People are grasping at straws in the noise.
Re: GPT-6 Astra, looped transformers, and hidden reasoning
#117Re: GPT-6 Astra, looped transformers, and hidden reasoning
#118I only used Astra while coding a bit so I can't comment on anything else but I have been really disappointed by it. It seems to overengineer really bad and it is also very slow due to it "thinking" too much I feel like. One example is that I asked it to implement a new functionality inside an existing App of mine and if I had written it myself it would have been like a ~50 line diff. Astra took like 10 minutes to wri…
Re: GPT-6 Astra, looped transformers, and hidden reasoning
#119What a clear and well-written article. I have only a basic understanding of LLM architecture and was able to follow along and gain intuition the whole time!
I got to > I am sure that OpenAI’s GPT-6 Astra is top of mind for everyone right now. and closed the tab.
Re: GPT-6 Astra, looped transformers, and hidden reasoning
#120Astra was insane until Monday but something happened on tuesday, now it feels like Sol. I grieve for the lost productivity but i hope they may give us the original Astra back.
Disagree completely. I started using Astra from Sol the day it was released, and was a virtually imperceptable difference and made lots of mistakes and shit architecture decisions from day 1 of release.
It prioritises getting something working over making something good during the 1-shot phase and outputs maximum slop.