Earlier quoted context omitted.
Do you have hard evidence of this assertion?
There is a toot from an Open AI person a couple days ago saying they are "pulling all the levers" because of capacity issues. I have no idea what the heck the person is talking about, but I'm guessing there are consequence for those levers. > "Demand for Astra is really unprecedented. We're pulling all the levers possible to sustain the demand, but I've not seen anything like it until now and we went through very ste…
GPT-6 Astra, looped transformers, and hidden reasoning
91–100 of 156 posts
Re: GPT-6 Astra, looped transformers, and hidden reasoning
#92Earlier quoted context omitted.
We can't have hard evidence. It's a SaaS and they own the code and the machine it runs on. So it may be a widespread hallucination. But there's no evidence of that either.
We could still have soft evidence though. Make a Todo app on Monday, and make a Todo app on Tuesday, and see what it makes in comparison.
Re: GPT-6 Astra, looped transformers, and hidden reasoning
#93Earlier quoted context omitted.
We can't have hard evidence. It's a SaaS and they own the code and the machine it runs on. So it may be a widespread hallucination. But there's no evidence of that either.
We could still have soft evidence though. Make a Todo app on Monday, and make a Todo app on Tuesday, and see what it makes in comparison.
Re: GPT-6 Astra, looped transformers, and hidden reasoning
#94I had not heard of looped transformers, but the engineering behind the number of loops per token / halting feels like trying to apply a diffusion process to a transformer while keeping the auto-regressive feature.
Re: GPT-6 Astra, looped transformers, and hidden reasoning
#95Earlier quoted context omitted.
We can't have hard evidence. It's a SaaS and they own the code and the machine it runs on. So it may be a widespread hallucination. But there's no evidence of that either.
We could still have soft evidence though. Make a Todo app on Monday, and make a Todo app on Tuesday, and see what it makes in comparison.
Re: GPT-6 Astra, looped transformers, and hidden reasoning
#96Earlier quoted context omitted.
Yeah q8 made so littler difference back when I was testing such things I'd be surprised if people could quickly notice that as a change. It's got to be either further quantized or some other type of optimization that kicks in when people notice the drop.
NVFP4 would buy them a huge increase in capacity but I think it would be noticeable.
Re: GPT-6 Astra, looped transformers, and hidden reasoning
#97Astra was insane until Monday but something happened on tuesday, now it feels like Sol. I grieve for the lost productivity but i hope they may give us the original Astra back.
Re: GPT-6 Astra, looped transformers, and hidden reasoning
#98Earlier quoted context omitted.
You think they introduce stronger quantization after a few days?
For sure they quickly move to q8, the output quality difference to bf16 is small compared to the speed/capacity gain
Re: GPT-6 Astra, looped transformers, and hidden reasoning
#99> So, the whole idea here is that we increase the effective depth from 22 to 44 block applications without adding another set of transformer weights. From what I gathered, LLM inference is bottlenecked on memory, right? Which implies there's "spare" compute we haven't been using? Does reusing the weights like this allow us to utilize it? (Do more math per unit of memory?)
Re: GPT-6 Astra, looped transformers, and hidden reasoning
#100Earlier quoted context omitted.
We could still have soft evidence though. Make a Todo app on Monday, and make a Todo app on Tuesday, and see what it makes in comparison.
someone already does that https://aistupidlevel.info/