Live data from Hacker News

Qwen3.8-Flash-Next

qwen.ai

241–246 of 246 posts

Re: Qwen3.8-Flash-Next

#241
post #58

> Qwen3.8-Flash-Next features a 125B-parameter main model, supplemented by an additional 51B N-gram embeddings, with 6B parameters activated per token. Didn’t see this mentioned yet. I wonder what this means for the effective size. It’s evidently ~176B paramètres, but how does that get quantized. A 4-bit quant under 100GB seems unlikely, I’m suspecting this won’t run in 128GB unified memory In principle I like the id…

I'm running the UD-Q4_K_S on my 128GB M5 Max with 180K context, it uses around 100GB~

Re: Qwen3.8-Flash-Next

#242

Earlier quoted context omitted.

> Given that OpenAI is ahead in intelligence, it's also reasonably likely that they are at the frontier of efficiency too. Frontier labs have no incentive to be at the frontier of efficiency. Claude still leads the pack in general intelligence yet has the worst efficiency by far.

They have an incentive to make their models efficient enough to serve demand and make a profit on it. The incentive that is missing is passing on efficiency improvements as price savings to customers, when your model is still in demand because of its higher intelligence.

Agreed, efficiency is still important, but being at the "frontier of efficiency" is significantly more relevant to commodity model providers than state-of-the-art model providers. Frontier labs are incentivized to route their spend towards beating benchmarks because that's what enables them to charge a premium.

Re: Qwen3.8-Flash-Next

#243
post #222

Earlier quoted context omitted.

... and once more with UD-IQ4_XS https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

how does it do with a pelican equipment case?

I guess they haven't benchmaxxed it yet?

https://imgur.com/a/vT636BS

Re: Qwen3.8-Flash-Next

#245
post #76
post #67

Earlier quoted context omitted.

Gonna have to wait a few days to see what the wizards of the HF community come up with…

They are already working on it. https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF https://unsloth.ai/docs/models/qwen3.8-next > You will need at least 75 GB of RAM or unified memory to run the model. Its smallest 1-bit quantized version is larger than usual because of the model’s architecture so 1-bit isn't really 1-bit at all. However, this also means the quantization is less aggressive, allowing the model to r…

It should be fine keeping the n-gram embeddings on SSD which lets you run at least the Q1 and Q2 models on 64GB

Re: Qwen3.8-Flash-Next

#246
post #126

I ran some pelicans at the four different reasoning levels (none, low, medium, xhigh - apparently high and xhigh are aliases of each other) on a DGX Spark using Unsloth's unsloth/Qwen3.8-Flash-Next-GGUF (UD-IQ1_S): https://tools.simonwillison.net/markdown-svg-renderer#url=ht... Surprised I didn't get one I liked as much as the Qwen 3.8 27B one https://simonwillison.net/2026/Aug/16/qwen-38-27b/#the-defau... , maybe be…

[dead]
Post reply on HN