Really happy for those with 128GB+ RAM. Sitting here with my Apple M1 Max with 64GB though. Was looking forward to a Qwen3.8-35B-A3B like many others.
Have you tested Muse Glimmer in low reasoning strength? Token generation is slow (and prefill is) but you will likely find it solves actual problems faster than Qwen 3.6 35B-A3B.
Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
41–50 of 178 posts
Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
#42Time to dust off my 128GB strix halo (literally—it’s been dusty, and it’s running a bit warm these days). Any idea where this model sits according toquality benchmarks? Pre-bubble MSRP on this hardware was $1400, and it draws 200-ish watts, putting it down into consumer territory. I’m wondering if it can replace claude for llm-friendly coding tasks.
So back in the Qwen 3.5 release, the 122B-A10B model scored slightly better than the 27B model. I'd expect this new 125B-A6B to perform similarly to the recently released 27B. Qwen3.8 27B is supposed to rival Sonnet/Opus 4.6.
4.6 ~= 4.8
4.7 much worse.
Fable and newer consistently tells me to pound sand, so I’m not sure what I’m paying $200/month for. 4.8 sometimes does too, but it’s at least usable most of the time.
So, I’d expect this to mostly replace Claude for my workflows. The main tradeoff for me should mostly be token throughput vs. no longer really trusting anthropic.
Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
#43Can I run a fp8 quant with 96gb VRAM?
Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
#44Earlier quoted context omitted.
You can run the 27B released last week. I haven't tried it yet myself but the 3.6 version runs great on my 5090.
Strongly recommend https://github.com/Neroued/ninfer , which can pull ~180 TPS on 5090 with 3.8, and 500 (!) with 3.6 35B-A3B.
Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
#45Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
#46Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
#47Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
#48Earlier quoted context omitted.
50-70tk/s is what I get on my m5 max on a 5-6bit Qwen 3.8 27B?
I don't know what black magic you're up to but I see more like 30-35t/s on a 16" M5 Max using 3.8:27b Q4, regardless of whether it's mlx or gguf. qwen3.5:122b-a10b is significantly faster at around 60-65.
Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
#49its worked out to to 40 tokens/seconds on their 80b-a3b model. we'll see how much of a hit this is.
Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
#50Really looking forward to this, 27B is a struggle with a strix halo and Laguna 2.1 can do stupid things for tooling calls.