2028 Could Bring the Most Mind-Bendingly Expensive Apple Product of All Time
1–4 of 4 posts
Re: 2028 Could Bring the Most Mind-Bendingly Expensive Apple Product of All Time
#2Re: 2028 Could Bring the Most Mind-Bendingly Expensive Apple Product of All Time
#3Are we going to get 1000 tps local?
In 2 years presumably the SOTA models will be much bigger, but I won’t be surprised if there’s a fable level sub 50B parameter model by that point either*, in which case 1000 tok/s may be reasonable.
* on artificial analysis intelligence index, GPT 4o is 12 and Qwen 35B A3B is 32.
Re: 2028 Could Bring the Most Mind-Bendingly Expensive Apple Product of All Time
#4Are we going to get 1000 tps local?
I’ve been using local models on a Framework desktop. I get > 50 tokens/s on Qwen 3.6 35B A3B, and I find that speed ok for coding (though the model not quite). If machine could do a SOTA model with say 0.5-2T parameters like GLM 5.2 at 100 tok/s, it would be very useable. In 2 years presumably the SOTA models will be much bigger, but I won’t be surprised if there’s a fable level sub 50B parameter model by that point…
Examples: Databases scale with data stored in them. It's a linear relationship. We have databases that load pages on demand as needed and keep hot ones in cache.
There is no 'database page' alternative for LLM's yet. Someday there will be and the paradigm shift will happen.
Local 100 tok/s+ is a gamechanger for sure.