Live data from Hacker News

Qwen3.8-Flash-Next

qwen.ai

21–30 of 246 posts

Re: Qwen3.8-Flash-Next

#21
Adding to my homelab stack, hopefully it doesn't overthink like the little model. Actually, hoping it thinks a bit less. Wait actually I'm really praying it reasons a bit more directly. But wait, I'm really sure that it must be a bit better.

Re: Qwen3.8-Flash-Next

#22
post #12

Didn't expect it to beat 3.8 27B so cleanly. Opus 4.6 Max self-hosted at 30 tok/s on a 5k Macbook in Aug 2026. The LLM timelines are crazy.

I don't like these comparisons. Sure it is impressive, but it does not have a world knowledge of larger models. It has most of theirs intelligence.

If/when we can get larger context this will mostly be mitigated by these smaller models being able to search the internet.

Self-learning/improving would be even better but that's still a long way to go.

Re: Qwen3.8-Flash-Next

#26
post #21

Adding to my homelab stack, hopefully it doesn't overthink like the little model. Actually, hoping it thinks a bit less. Wait actually I'm really praying it reasons a bit more directly. But wait, I'm really sure that it must be a bit better.

That's low reasoning for a model, but max for a HN comment.

Re: Qwen3.8-Flash-Next

#27
post #12

Didn't expect it to beat 3.8 27B so cleanly. Opus 4.6 Max self-hosted at 30 tok/s on a 5k Macbook in Aug 2026. The LLM timelines are crazy.

I don't like these comparisons. Sure it is impressive, but it does not have a world knowledge of larger models. It has most of theirs intelligence.

For world knowledge, you'd want it to find and reference the source material to be sure. At that point, it doesn't matter if the knowledge is embedded.

Re: Qwen3.8-Flash-Next

#28
post #15

Didn't expect it to beat 3.8 27B so cleanly. Opus 4.6 Max self-hosted at 30 tok/s on a 5k Macbook in Aug 2026. The LLM timelines are crazy.

>Opus 4.6 Max self-hosted at 30 tok/s on a 5k Macbook in Aug 2026. The LLM timelines are crazy. How much memory does this translate to and what quantization (if any) were applied?

128GB, 4-bit quantized.

Re: Qwen3.8-Flash-Next

#30
post #25

It's in Unsloth Desktop already. Looks like it's 73GB, so 128GB Mac or Strix Halo etc will work. Exciting!

Download is available, but likely need to wait for an update, I get this which is understandable with the architectural change :

Original error: llama.cpp does not support this GGUF's model architecture ('qwen4exp')

Edit : Saw the pull request, should arrive soon enough https://github.com/ggml-org/llama.cpp/pull/27742

Post reply on HN