Live data from Hacker News

Qwen3.8-Flash-Next

qwen.ai

231–240 of 246 posts

Re: Qwen3.8-Flash-Next

#231

Earlier quoted context omitted.

the important thing is that Qwen 3.7 27B will run unlimited jobs on my consumer grade laptop at 60 tokens/second for free, forever, in about 1-2 years

Thats only important if running it locally is critical for privacy reasons or just as a hobby. Time has a cost in business. If a model needs 30 million tokens to achieve a similar result as another that can do it in 10 million, that 60 tokens per second will take a long time.

Right now qwen 3.6 35b-a3b has a success rate of 92% and qwen 3.8 27b has a success rate of 96%. But the 35b moe does about 1080 tokens/s at concurrency 54, vs 480 tokens/s at concurrency 28. For our specific workflow on blackwell.

Of course enormous batch jobs are different. I was explicit when I said consumer laptop.

Re: Qwen3.8-Flash-Next

#232
post #86

Earlier quoted context omitted.

Those prices are just tokens? Since each model uses different amounts of tokens to do the same thing, it's a misleading price that often makes open-weights look more competitive than they are, since most open weights models use dramatically more tokens and time to complete tasks than many frontier models. In Artifical Analysis's cost per task, Luna(max) costs $0.05 per task, and Qwen 3.8 27B costs $0.25 per task, a 5…

the important thing is that Qwen 3.7 27B will run unlimited jobs on my consumer grade laptop at 60 tokens/second for free, forever, in about 1-2 years

It's not free. You're paying electricity and you're ignoring the cost of the hardware. Even on electricity alone, there are cloud providers who may beat your laptop on price per million tokens. Qwen 3.8 flash is interesting in this space.

Not to say that there aren't other benefits of running models locally, I loaded Qwen 3.8 27B 6bit MLX just yesterday.

Re: Qwen3.8-Flash-Next

#233
post #214

Earlier quoted context omitted.

Why would anyone car what the launch price is? Comparing launch pricing is just an odd thing to do.

Because labs can learn to optimize inference post launch, plus can move to use bigger/better clusters depending on demand. It is not impossible to imagine Qwen cuts prices further with QAT/MTP-like improvements.

Or they could move from highly subsidized models like the Deepseek 4 launch pricing.

Launch price is just like any other price. It's just a price. It's impossible to guess what might or might not happen.

Compare the price now.

Re: Qwen3.8-Flash-Next

#235

Earlier quoted context omitted.

I find LLMisms very annoying to read, it’s almost like they are bullet points in the shape of a paragraph. It feels very “skippy” to me. EDITED: Removed a question that I couldn’t make feel suitably polite.

Suppose you time-zap a modern physics curriculum on a solarpowered computer tablet to any shortly-pre-Galilean era and observe their reaction to the course notes. In that era, plenty of fields required mathematics, engineering and architecture. The church would prescribe and uphold Aristotelean Logic "When objects fall, they fall down" style statements (never mind that if you throw an object up, it doesn't instantly…

Good writing is generally writing that communicates the intended meaning. Transmitting thought and meaning is inherently lossy and the content is irrelevant if it is insoluble in the mind of the recipient.

LLMs aren’t really great at this yet and I think the solution is, hopefully, that they improve. Anything else is accommodating a tool that should be accommodating the user.

Re: Qwen3.8-Flash-Next

#236
post #58

> Qwen3.8-Flash-Next features a 125B-parameter main model, supplemented by an additional 51B N-gram embeddings, with 6B parameters activated per token. Didn’t see this mentioned yet. I wonder what this means for the effective size. It’s evidently ~176B paramètres, but how does that get quantized. A 4-bit quant under 100GB seems unlikely, I’m suspecting this won’t run in 128GB unified memory In principle I like the id…

People in my server are running it on Strix Halo 128GB using RoCmFP4 and reporting 35tok/s, without much optimization, with proper MTP, better kernel, expecting about 50-60tok/s.

The largest unsloth published gguf also fits and runs just fine on a cpu-only machine with 128GB RAM, using llama-server PR 27742

https://github.com/ggml-org/llama.cpp/pull/27742

Re: Qwen3.8-Flash-Next

#237
post #126

I ran some pelicans at the four different reasoning levels (none, low, medium, xhigh - apparently high and xhigh are aliases of each other) on a DGX Spark using Unsloth's unsloth/Qwen3.8-Flash-Next-GGUF (UD-IQ1_S): https://tools.simonwillison.net/markdown-svg-renderer#url=ht... Surprised I didn't get one I liked as much as the Qwen 3.8 27B one https://simonwillison.net/2026/Aug/16/qwen-38-27b/#the-defau... , maybe be…

If I read correctly, that's based on a 1-bit quantization, and can we really expect that to produce any useful output at all?

if the goal is vector art salvador dali, like melting clocks and stuff, sure

Re: Qwen3.8-Flash-Next

#238
post #222
post #173

Earlier quoted context omitted.

Tried again with a different quant, UD-Q2_K_XL: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

... and once more with UD-IQ4_XS https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

how does it do with a pelican equipment case?

Re: Qwen3.8-Flash-Next

#239
post #49

Earlier quoted context omitted.

I only see a 1-bit quant posted on unsloth HF and it’s 72.5 GB. Is that what you mean? That’s much bigger than I expected. If you can’t run a 4 bit quant in on Strix Halo it becomes a lot less interesting. https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF

In their page they say it will need at least 112GB[0], so including context, that would be a tight fit. I'm also hoping I can make a q4 fit on my 128GB strix halo [0]: https://unsloth.ai/docs/models/qwen3.8-next#qwen3.8-flash-ne...

in llama-server PR 27742 it fits fine in 128GB RAM on a CPU only system , this is with --load-mode mlock to stuff the whole thing persistently into memory at llama-server launch time, no mmap

0.01.033.250 I common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted |

0.01.033.261 I common_memory_breakdown_print: | - Host | 118186 = 106166 + 8898 + 3122 |

0.01.092.684 I common_params_fit_impl: projected to use 118186 MiB of host memory vs. 128855 MiB of total host memory

Re: Qwen3.8-Flash-Next

#240
post #201

Earlier quoted context omitted.

Quoting RGFusion from Reddit: LLMs run into an issue where the further you train a model, the more it overwrites facts with generalized concepts. You need the model to be able to do both. Intelligence arises from generalization, but without accurate information the model will hallucinate. The engram table allows for a low-computational method of fact-recall. You can think of it like a better form of RAG, where the da…

Very interesting. Is this compatible with MoE architectures? I'm not too familiar with how this works.

By the way; I found that DeepSeek says that it is even an "ideal complement to modern MoE architecture"

https://arxiv.org/abs/2601.07372v1

Post reply on HN