Live data from Hacker News

Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

modelscope.cn

71–80 of 178 posts

Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

#71

Really looking forward to this, 27B is a struggle with a strix halo and Laguna 2.1 can do stupid things for tooling calls.

Yeah 27B is way too slow for the Strix Halo. Laguna was better but still slow when I tried it. Qwen3.6 35B is still the best today.

What sort of pp/tg speed do you get on a Strix Halo?

Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

#73
post #65

Earlier quoted context omitted.

That’s only true if you think AI is the only reason to own a powerful and efficient server. Mine does plenty of traditional server stuff too.

I can do traditional server stuff on any old computer with a big hard disk and a decent amount of RAM. That's not worth $3500-$4000. When RAMpocalypse is over and we can buy a Strix Halo for under $2000 again, the math starts mathing. It becomes a pretty great desktop computer that also happens to run AI pretty well at a pretty good price.

Yeah, but that computer can’t also do the AI stuff. And not everybody has a desktop with multiple 32GB GPUs available.

I’ll admit though I’m biased because I bought my board for $1600 back before the prices went crazy.

Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

#74

Very curious to see how this compares to Deepseek v4 Flash. I have to assume they wouldn't be releasing this if it was worse.

Their "next" variants are usually undercooked, but useful for the community to verify support for inference stacks. This will likely be the same.

Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

#76
post #2

Can you share the source for the parameter count (125B A6B)? I didn't see it anywhere in the page.

This is what I copied from the en version of the modelscope page, right when they published it:

> Redisgned Multimodal MoE Model: 125B main model parameters, supplemented by an additional 51B N-gram embeddings,and 6B parameters activated per token.

> Efficient Training and Inference: Significantly reduces training and inference costs. At ~1/9th the training cost,Qwen3.8-Flash-Next achieves comparable capability against Qwen3.7-Plus, while being more capable in areas of coding and cowork.

There was another paragraph about a new attention, but I didn't copy that.

Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

#79

Really looking forward to this, 27B is a struggle with a strix halo and Laguna 2.1 can do stupid things for tooling calls.

Yeah 27B is way too slow for the Strix Halo. Laguna was better but still slow when I tried it. Qwen3.6 35B is still the best today.

Have you given Ornith-1.5-35B a shot?

It's been a pretty decent step up for me compared to Qwen3.6

https://news.ycombinator.com/item?id=49362401

Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

#80

Earlier quoted context omitted.

Yeah 27B is way too slow for the Strix Halo. Laguna was better but still slow when I tried it. Qwen3.6 35B is still the best today.

What sort of pp/tg speed do you get on a Strix Halo?

This is the best I got, all with Unsloth's quantizations.

Laguna-S-2.1:UD-Q4_K_XL (no MTP) pp=186.4 t/s tg=27.8 t/s

Qwen3.6-35B:UD-Q4_K_XL (with MTP) pp=404.4 t/s tg=83.2 t/s

Qwen3.6-27B:UD-Q4_K_XL (recorded pre-MTP) pp=343 t/s tg=12.1 t/s

Laguna actually performed better than I remembered. I thought it was slower.

Post reply on HN