I've been saying this for a couple months now since I got decent hardware and started using my local Qwen 3.6 exclusively. I have no doubt the future for individuals and medium-sized companies is local private AI.
Could you share some of your hardware details for Qwen 3.6? And are you using the dense or MoE variant?
That is pretty usable. You could get 65t/s or more with MTP, but only if you drop the context size, which I would advise against.
Results are better with 256k context and a larger quant, however, that's not going to fit on the 4090 you already had lying around for playing cyberpunk 2077.
The MoE models make me rather unhappy. Idk. They feel braindead to me, but YMMV.