Earlier quoted context omitted.
Their previous model Qwen3.5 was available in many sizes, from very small sizes intended for smartphones, to medium sizes like 27B and big sizes like 122B and 397B. This model is the first that is provided with open weights from their newer family of models Qwen3.6. Judging from its medium size, Qwen/Qwen3.6-35B-A3B is intended as a superior replacement of Qwen/Qwen3.5-27B. It remains to be seen whether they will als…
> Qwen/Qwen3.6-35B-A3B is intended as a superior replacement of Qwen/Qwen3.5-27B Not at all, Qwen3.5-27B was much better than Qwen3.5-35B-A3B (dense vs MoE).
Qwen3.6-35B-A3B: Agentic coding power, now open to all
361–370 of 563 posts
Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#362Honestly, this is the AI software I actually look forward to seeing. No hype about it being too dangerous to release. No IPO pumping hype. No subscription fees. I am so pumped to try this!
Same here. I really hope in a near future local model will be good enough and hardware fast enough to run them to become viable for most use cases
Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#363I've been running this on my laptop with the Unsloth 20.9GB GGUF in LM Studio: https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF/blob/mai... It drew a better pelican riding a bicycle than Opus 4.7 did! https://simonwillison.net/2026/Apr/16/qwen-beats-opus/
The flamingo on Qwen's unicycle is sitting on the tire, not the seat. That wins because of sunglasses?
Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#364What is the min VRAM this can run on given it is MOE?
Fwiw, with its predecessor's Qwen3.5-35B-A3B-Q6_K.gguf, on a laptop's 6 GB VRAM and 32 GB RAM, with default llama.cpp settings, I get 20 t/s generation.
1 - https://github.com/ggml-org/llama.cpp/blob/master/docs/build...
Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#365Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#366Earlier quoted context omitted.
Fwiw, with its predecessor's Qwen3.5-35B-A3B-Q6_K.gguf, on a laptop's 6 GB VRAM and 32 GB RAM, with default llama.cpp settings, I get 20 t/s generation.
Have you tried running llama.cpp with Unified Memory Access[1] so your iGPU can seamlessly grab some of the RAM? The environment variable is prefixed with CUDA but this is not CUDA specific. It made a pretty significant difference (> 40% tg/s) on my Ryzen 7840U laptop. 1 - https://github.com/ggml-org/llama.cpp/blob/master/docs/build...
Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#367Earlier quoted context omitted.
The 27B model is dense. Releasing a dense model first would be terrible marketing, whereas 35A3B is a lot smarter and more quick-witted by comparison!
What? 35B-A3B is not nearly as smart as 27B.
Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#368Earlier quoted context omitted.
Yep we can do that probs add a table - in general be post in discussions of model pages - for eg https://huggingface.co/unsloth/MiniMax-M2.7-GGUF/discussions... HF also provides SHA256 for eg https://huggingface.co/unsloth/MiniMax-M2.7-GGUF/blob/main/U... is 92986e39a0c0b5f12c2c9b6a811dad59e3317caaf1b7ad5c7f0d7d12abc4a6e8 But agreed it's probs better to place them in a table
Thanks! I know about HF's chunk checksums, but HF doesn't publish (or possibly even know) the merged checksums.
Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#369Earlier quoted context omitted.
Why are you looking to move off Ollama? Just curious because I'm using Ollama and the cloud models (Kimi 2.5 and Minimax 2.7) which I'm having lots of good success with.
Ollama co mingles online and local models which defeats the purpose for me
Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#370Earlier quoted context omitted.
There are way too many good uses of these models for local that I fully expect a standard workstation 10 years from now to start at 128GB of RAM and have at least a workstation inference device.
or if you believe a lot of HN crowd we are in AI bubble and in 10 years inference will be dirt cheap when all of this crashes and we have all this hardware in data centers and it won't make any sense to run monster workstations at home (I work 128GB M4 but not run inference, just too many electron apps running at the same time...) :)