Live data from Hacker News

Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

modelscope.cn

41–50 of 178 posts

Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

#41
post #23

Really happy for those with 128GB+ RAM. Sitting here with my Apple M1 Max with 64GB though. Was looking forward to a Qwen3.8-35B-A3B like many others.

Have you tested Muse Glimmer in low reasoning strength? Token generation is slow (and prefill is) but you will likely find it solves actual problems faster than Qwen 3.6 35B-A3B.

I’ll give it a try! Thanks for the heads up!

Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

#42
post #26

Time to dust off my 128GB strix halo (literally—it’s been dusty, and it’s running a bit warm these days). Any idea where this model sits according toquality benchmarks? Pre-bubble MSRP on this hardware was $1400, and it draws 200-ish watts, putting it down into consumer territory. I’m wondering if it can replace claude for llm-friendly coding tasks.

So back in the Qwen 3.5 release, the 122B-A10B model scored slightly better than the 27B model. I'd expect this new 125B-A6B to perform similarly to the recently released 27B. Qwen3.8 27B is supposed to rival Sonnet/Opus 4.6.

Thanks. My current stack ranking of anthropic models is:

4.6 ~= 4.8

4.7 much worse.

Fable and newer consistently tells me to pound sand, so I’m not sure what I’m paying $200/month for. 4.8 sometimes does too, but it’s at least usable most of the time.

So, I’d expect this to mostly replace Claude for my workflows. The main tradeoff for me should mostly be token throughput vs. no longer really trusting anthropic.

Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

#44
post #35

Earlier quoted context omitted.

You can run the 27B released last week. I haven't tried it yet myself but the 3.6 version runs great on my 5090.

Strongly recommend https://github.com/Neroued/ninfer , which can pull ~180 TPS on 5090 with 3.8, and 500 (!) with 3.6 35B-A3B.

I've been waiting for the dust to settle on this model so I can find a good runtime setup. I'm definitely bookmarking this. Thanks!

Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

#45
post #20

Earlier quoted context omitted.

I have a 128GB M5 Max, and it sucks at this stage. 50-70 tok/s might be something...

50-70tk/s is what I get on my m5 max on a 5-6bit Qwen 3.8 27B?

I tried 8-bit, perhaps I should try 6-bit.

Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

#48
post #38
post #20

Earlier quoted context omitted.

50-70tk/s is what I get on my m5 max on a 5-6bit Qwen 3.8 27B?

I don't know what black magic you're up to but I see more like 30-35t/s on a 16" M5 Max using 3.8:27b Q4, regardless of whether it's mlx or gguf. qwen3.5:122b-a10b is significantly faster at around 60-65.

With MTP? I get 25-30 TPS on a strix halo. 50+ on a M5 max should very doable. Dflash (2) will push your TG even further

Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

#49
This is great. I have a weird system layout (192gb system ram, 8gb vram). the mixture of experts models have been nice when i can run the dense reasoning layers on the gpu (which somehow fit?!) and then the expert on the cpu.

its worked out to to 40 tokens/seconds on their 80b-a3b model. we'll see how much of a hit this is.

Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

#50

Really looking forward to this, 27B is a struggle with a strix halo and Laguna 2.1 can do stupid things for tooling calls.

Yeah 27B is way too slow for the Strix Halo. Laguna was better but still slow when I tried it. Qwen3.6 35B is still the best today.
Post reply on HN