> Qwen3.8-Flash-Next features a 125B-parameter main model, supplemented by an additional 51B N-gram embeddings, with 6B parameters activated per token. Didn’t see this mentioned yet. I wonder what this means for the effective size. It’s evidently ~176B paramètres, but how does that get quantized. A 4-bit quant under 100GB seems unlikely, I’m suspecting this won’t run in 128GB unified memory In principle I like the id…
Qwen3.8-Flash-Next
241–246 of 246 posts
Re: Qwen3.8-Flash-Next
#242Earlier quoted context omitted.
> Given that OpenAI is ahead in intelligence, it's also reasonably likely that they are at the frontier of efficiency too. Frontier labs have no incentive to be at the frontier of efficiency. Claude still leads the pack in general intelligence yet has the worst efficiency by far.
They have an incentive to make their models efficient enough to serve demand and make a profit on it. The incentive that is missing is passing on efficiency improvements as price savings to customers, when your model is still in demand because of its higher intelligence.
Re: Qwen3.8-Flash-Next
#243Earlier quoted context omitted.
... and once more with UD-IQ4_XS https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
how does it do with a pelican equipment case?
Re: Qwen3.8-Flash-Next
#244Re: Qwen3.8-Flash-Next
#245Earlier quoted context omitted.
Gonna have to wait a few days to see what the wizards of the HF community come up with…
They are already working on it. https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF https://unsloth.ai/docs/models/qwen3.8-next > You will need at least 75 GB of RAM or unified memory to run the model. Its smallest 1-bit quantized version is larger than usual because of the model’s architecture so 1-bit isn't really 1-bit at all. However, this also means the quantization is less aggressive, allowing the model to r…
Re: Qwen3.8-Flash-Next
#246I ran some pelicans at the four different reasoning levels (none, low, medium, xhigh - apparently high and xhigh are aliases of each other) on a DGX Spark using Unsloth's unsloth/Qwen3.8-Flash-Next-GGUF (UD-IQ1_S): https://tools.simonwillison.net/markdown-svg-renderer#url=ht... Surprised I didn't get one I liked as much as the Qwen 3.8 27B one https://simonwillison.net/2026/Aug/16/qwen-38-27b/#the-defau... , maybe be…