Really looking forward to this, 27B is a struggle with a strix halo and Laguna 2.1 can do stupid things for tooling calls.
Yeah 27B is way too slow for the Strix Halo. Laguna was better but still slow when I tried it. Qwen3.6 35B is still the best today.
Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
71–80 of 178 posts
Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
#72+1 to the long list of people hoping for Qwen3.8-27b A3B.
Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
#73Earlier quoted context omitted.
That’s only true if you think AI is the only reason to own a powerful and efficient server. Mine does plenty of traditional server stuff too.
I can do traditional server stuff on any old computer with a big hard disk and a decent amount of RAM. That's not worth $3500-$4000. When RAMpocalypse is over and we can buy a Strix Halo for under $2000 again, the math starts mathing. It becomes a pretty great desktop computer that also happens to run AI pretty well at a pretty good price.
I’ll admit though I’m biased because I bought my board for $1600 back before the prices went crazy.
Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
#74Very curious to see how this compares to Deepseek v4 Flash. I have to assume they wouldn't be releasing this if it was worse.
Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
#75gpt oss killer? this can easily run on a server cpu with its memory bandwidth
Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
#76Can you share the source for the parameter count (125B A6B)? I didn't see it anywhere in the page.
> Redisgned Multimodal MoE Model: 125B main model parameters, supplemented by an additional 51B N-gram embeddings,and 6B parameters activated per token.
> Efficient Training and Inference: Significantly reduces training and inference costs. At ~1/9th the training cost,Qwen3.8-Flash-Next achieves comparable capability against Qwen3.7-Plus, while being more capable in areas of coding and cowork.
There was another paragraph about a new attention, but I didn't copy that.
Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
#77gpt oss killer? this can easily run on a server cpu with its memory bandwidth
Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
#78Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
#79Really looking forward to this, 27B is a struggle with a strix halo and Laguna 2.1 can do stupid things for tooling calls.
Yeah 27B is way too slow for the Strix Halo. Laguna was better but still slow when I tried it. Qwen3.6 35B is still the best today.
It's been a pretty decent step up for me compared to Qwen3.6
Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
#80Earlier quoted context omitted.
Yeah 27B is way too slow for the Strix Halo. Laguna was better but still slow when I tried it. Qwen3.6 35B is still the best today.
What sort of pp/tg speed do you get on a Strix Halo?
Laguna-S-2.1:UD-Q4_K_XL (no MTP) pp=186.4 t/s tg=27.8 t/s
Qwen3.6-35B:UD-Q4_K_XL (with MTP) pp=404.4 t/s tg=83.2 t/s
Qwen3.6-27B:UD-Q4_K_XL (recorded pre-MTP) pp=343 t/s tg=12.1 t/s
Laguna actually performed better than I remembered. I thought it was slower.