Live data from Hacker News

Qwen3.8-Flash-Next

qwen.ai

11–20 of 246 posts

Re: Qwen3.8-Flash-Next

#11
post #4

do we really need breaking news about qwen posted every single day?

I and presumably quite a few others with AMD AI or Apple Mac platforms are very impacted by this.

:)

It is very relevant and for a certain group of us, far more impactful to our work the next month(s) than any blog post could be.

Re: Qwen3.8-Flash-Next

#12

Didn't expect it to beat 3.8 27B so cleanly. Opus 4.6 Max self-hosted at 30 tok/s on a 5k Macbook in Aug 2026. The LLM timelines are crazy.

I don't like these comparisons. Sure it is impressive, but it does not have a world knowledge of larger models. It has most of theirs intelligence.

Re: Qwen3.8-Flash-Next

#13
post #4

do we really need breaking news about qwen posted every single day?

If there’s news, then yes. This is a pretty great new release for those still stuck on Qwen3.6 35B A3B if they have enough memory but don’t have super powerful compute.

I wonder if I could get this running through vLLM on 6x Nvidia L4 - the 3.6 worked great on 4 cards but sadly TP6 just isn’t a thing and I don’t have 8 cards available, maybe it’s gonna be okay with like TP2 and MTP. I have no idea at this time, probably need to test out what even might be possible.

Re: Qwen3.8-Flash-Next

#15

Didn't expect it to beat 3.8 27B so cleanly. Opus 4.6 Max self-hosted at 30 tok/s on a 5k Macbook in Aug 2026. The LLM timelines are crazy.

>Opus 4.6 Max self-hosted at 30 tok/s on a 5k Macbook in Aug 2026. The LLM timelines are crazy.

How much memory does this translate to and what quantization (if any) were applied?

Re: Qwen3.8-Flash-Next

#17

Didn't expect it to beat 3.8 27B so cleanly. Opus 4.6 Max self-hosted at 30 tok/s on a 5k Macbook in Aug 2026. The LLM timelines are crazy.

For comparison with hosted models, GPT 5.6 Luna scores 67% on DeepSWE, compared to 59% here for Qwen.

Luna is $0.20 / $1.20 vs $0.16 / $0.47 with Qwen.

Re: Qwen3.8-Flash-Next

#18
post #4

do we really need breaking news about qwen posted every single day?

This particular release is interesting because it's a preview of qwen4 architecture. And, while benchmarks are iffy, this is a direct comparison, by the same team, with qwen3.8-27b that was pretty well received for a local model.

This "next" release adds a new concept, first public release with n-grams, I think. And it's in a MoE size that is likely to be very fast and cheap to serve (faster than 27b for sure). It's also well suited for inference on alternative compute (i.e. sparks, macs, etc) so it's relevant to local users.

Re: Qwen3.8-Flash-Next

#19
post #4

do we really need breaking news about qwen posted every single day?

This actually is meaningful news, I think. Pretty wide audience appeal in the local LLM space too.

Re: Qwen3.8-Flash-Next

#20
post #4

do we really need breaking news about qwen posted every single day?

There are many topics, personalities and politicians we hear about daily who have no merit.

Qwen's advances do (currently) have merit.

Post reply on HN