Live data from Hacker News

Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

modelscope.cn

121–130 of 178 posts

Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

#121

Finally, a reason to own a 128GB Strix Halo or GB10 device. Or a reason to consider the new Mac Studio. I have a Strix Halo and dual 32GB GPUs in my desktop, and the latter is pretty much always better for running local models because it's quite a bit faster due to higher memory bandwidth. There simply haven't been any models that are better than Qwen 27B or Gemma 31B, which run comfortably in 64GB with big context.…

I have had an impossible time getting 120B or better models running on Strix Halo (especially under Windows) with any large context windows. And 30-40 tokens/second is fine, but not the fastest.

For the most part lately I have been sticking with Qwen 3.8 27b and that thing will easily suck up 64gb of ram. Add in docker with some additional programs running and it's really easy to eat up 128gb of ram.

Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

#122

Alibaba is giving sleepless nights to the tech giants

To be fair, Alibaba IS a tech giant, one of the biggest in fact. They are just giving sleepless nights to the western tech giants.

I'm hearing Tencent, Zhipu and Baidu shaking from here. It's fair to assume BATX / 6 Tigers don't sleep very well either.

Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

#123

Earlier quoted context omitted.

A 3060ti 8gb, released in 2020, has 448 GB/s of bandwidth compared to the Halo 256 GB/s The 3080ti is 912.4 GB/s

And the newly announced/launched Apple M6 has 170GB/s of unified memory bandwidth, meanwhile M5 Ultra gets 1.2TB/s of unified memory bandwidth. https://www.apple.com/newsroom/2026/08/apple-introduces-m6-a... Not sure if the first one is a typo on their press release, can't be just 170GB/s then be pushed for AI use, can it? Could be a different measurement I suppose...

You got me curious so I looked up the previous chips[0].

Memory bandwidth M1: 68 GB/s M2: 100 GB/s (47% increase) M3: 100 GB/s (0% increase) M4: 120 GB/s (20% increase) M5: 153 GB/s (27.5% increase)

So, M6: 170 GB/s (11% increase) doesn’t seem impossible, though I would have expected more.

[0]: https://www.jdhodges.com/blog/apple-cpu-compared-m1-m3-m3-m4...

Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

#124

Finally, a reason to own a 128GB Strix Halo or GB10 device. Or a reason to consider the new Mac Studio. I have a Strix Halo and dual 32GB GPUs in my desktop, and the latter is pretty much always better for running local models because it's quite a bit faster due to higher memory bandwidth. There simply haven't been any models that are better than Qwen 27B or Gemma 31B, which run comfortably in 64GB with big context.…

I have had an impossible time getting 120B or better models running on Strix Halo (especially under Windows) with any large context windows. And 30-40 tokens/second is fine, but not the fastest. For the most part lately I have been sticking with Qwen 3.8 27b and that thing will easily suck up 64gb of ram. Add in docker with some additional programs running and it's really easy to eat up 128gb of ram.

I found a couple of different 4-bit quantizations of Laguna S 2.1 that run pretty well with pretty big context (also quantized, to 8 bits, I think). Unfortunately, Laguna isn't better than Qwen 3.8 27B, which I'm able to run at roughly the same speed on my desktop machine, so I don't use Laguna or the Strix Halo very much, lately. (It's also too hot for me to be running heaters for inference. It's been ~110F most days for the past few weeks.)

Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

#125

Earlier quoted context omitted.

They've said no moe for 3.8, and since they're already releasing a qwen4 early preview, they're probably focusing on that arch going forward.

Where did they say that? My understanding of this 3.8-Flash-Next release is that it's a MOE (as per the title of the posting here, 125B a6b)

A bit of context: 3.5 was the last version where they released their entire suite of models 2b-400b. Then 3.6 got a 27b dense and a 35b moe. Then 3.7 was API only, and 3.8 got only the 27b dense. The devs confirmed on twitter that 35b moe would not come. So that's what I meant by 3.8 is not getting a moe.

3.8 next is not really a 3.8 (but I guess they had to disambiguate from the previous next). It's a preview of qwen4 architecture (and it is an moe + ngram), released early as a preview, and to help the community sort out inference before qwen4 releases.

Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

#126

Earlier quoted context omitted.

So back in the Qwen 3.5 release, the 122B-A10B model scored slightly better than the 27B model. I'd expect this new 125B-A6B to perform similarly to the recently released 27B. Qwen3.8 27B is supposed to rival Sonnet/Opus 4.6.

Qwen3.8/Qwen3.6 has a weird self doubt/thinking too much problem. You can prompt it away. I would say it "approximates" Opus 4.X class models well enough especially for coding/linux problems. The only reason I stopped using it as much is I was getting 25-35tok/s on Intel B70 (non-quant) which made some responses slow. For a long running/autonomous task, it would probably be sufficient.

Two things:

- check your temp settings vs. Qwen's recommendations, they specify what it should be in the model card for thinking on, and that reduces some "over" think.

- the model appears to be intentionally designed to do a lot more test-time compute, if you anthropomorphize the tokens, it looks like overthinking and anxiety, but it is just spending compute to get to the end result, so it may not actually help it to prompt down the token spend (depending on the problem)

Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

#127

Earlier quoted context omitted.

A 3060ti 8gb, released in 2020, has 448 GB/s of bandwidth compared to the Halo 256 GB/s The 3080ti is 912.4 GB/s

And the newly announced/launched Apple M6 has 170GB/s of unified memory bandwidth, meanwhile M5 Ultra gets 1.2TB/s of unified memory bandwidth. https://www.apple.com/newsroom/2026/08/apple-introduces-m6-a... Not sure if the first one is a typo on their press release, can't be just 170GB/s then be pushed for AI use, can it? Could be a different measurement I suppose...

From a bandwidth perspective, the ultra is like 8 M5s fused together (@ 150GB/s), that's how it gets to the 1200.

Historically the Pro doubles the base, the Max doubles the Pro, and the Ultra doubles the Max.

If an M6 ultra were released today it would be 1.36TB/s.

Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

#128

Earlier quoted context omitted.

You can already run it locally its just not the same. It is still slow, a lot slower than what you are used to with claude and co. And as soon as you increase context size, your memory requirements jump. Then when it runs for 30 minutes for something claude needs 5, your device will get hot. And even a used 3090 is apparently now between 1-2k.

> It is still slow, a lot slower than what you are used to with claude and co. That really depends on the model, I run a few models locally. All at speeds comparable to or faster than Opus. In general we haven't reached the ceiling for what performance we can get out of consumer hardware. As evidence by FreeToken which hasn't even added MTP/speculative drafting support yet, which will add another boost. > Then when i…

I have 2 4090 and last time i played around with it, the context window killed it for me.

The normal LLMStudio stuff works great, but then i tried out anything with subagent things or parallel stuff and it trashed my cache and got super sluggish/slowish.

What do you run and how?

Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

#129
post #107

Earlier quoted context omitted.

> It is still slow, a lot slower than what you are used to with claude and co. That really depends on the model, I run a few models locally. All at speeds comparable to or faster than Opus. In general we haven't reached the ceiling for what performance we can get out of consumer hardware. As evidence by FreeToken which hasn't even added MTP/speculative drafting support yet, which will add another boost. > Then when i…

Which models when run locally come close to Sol and Opus, from your experience? And which harness do you use?

Qwen 3.8 27B is so far the closest I have run, though it suffers on speed compared to Gemma4 26B MoE model which I still use. Neither are going to match Opus or Sol though, but they can be as fast or faster depending on what you are using them for.

I don't use a harness, all the mainstream ones I have tried have tanked my productivity. I know that's not a common sentiment, but it's been my experience. None seem built for the way I work. I program mostly in my head first away from keyboard, then go type it out (faster than it would take to describe the solution to an LLM). Also a perfectionist who likes to learn, and tends to work on out of distribution problems. All I need is a simple chat interface for light research, quick small scoped prototypes, and generating simple scripts.

Not saying a harness is out of the question for me, just all I have seen and tested so far are not for me. Maybe if someone builds a more deterministic harness that doesn't rely on plain English skill files that bloat context and only sometimes do what you want.

Re: Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

#130
post #38

Earlier quoted context omitted.

I don't know what black magic you're up to but I see more like 30-35t/s on a 16" M5 Max using 3.8:27b Q4, regardless of whether it's mlx or gguf. qwen3.5:122b-a10b is significantly faster at around 60-65.

With MTP? I get 25-30 TPS on a strix halo. 50+ on a M5 max should very doable. Dflash (2) will push your TG even further

It's a bit deceptive to state inference speeds without mentioning the additional things you're doing to achieve them
Post reply on HN