Earlier quoted context omitted.
yeah the 27B feels like something completely different. If you use it on long context tasks it performs WAY better than 35b-a3b
I've been telling analysts/investors for a long time that dense architectures aren't "worse" than sparse MoEs and to continue to anticipate the see-saw of releases on those two sub-architectures. Glad to continuously be vindicated on this one. For those who don't believe me. Go take a look at the logprobs of a MoE model and a dense model and let me know if you can notice anything. Researchers sure did.
Qwen3.6-35B-A3B: Agentic coding power, now open to all
451–460 of 563 posts
Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#452Qwen3.6 and Gemma4 have the same issue of never getting to the point and just getting stuck in never ending repeating thought loops. Qwen3.5 is still the best local model that works.
Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#453Earlier quoted context omitted.
Where did you see a haiku comparison? Haiku 4.5 was my daily driver for a month or so before Opus 4.5 dropped and would be unreasonably happy if a local model can give me similar capability
Artificial Analysis hasn't posted their independent analysis of Qwen3.6 35B A3B yet, but Alibaba's benchmarks paint it as being on par with Qwen3.5 27B (or better in some cases). Even Qwen3.5 35B A3B benchmarks roughly on par with Haiku 4.5, so Qwen3.6 should be a noticeable step up. https://artificialanalysis.ai/models?models=gpt-oss-120b%2Cg... No, these benchmarks are not perfect, but short of trying it yourself,…
Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#454I'm broadly curious how people are using these local models. Literally, how are they attaching harnesses to this and finding more value than just renting tokens from Anthropic of OpenAI?
Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#455Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#456Earlier quoted context omitted.
While they can be run locally, and most of the discussion on HN about that, I bet that if you look at total tok/day local usage is a tiny amount compared to total cloud inference even for these models. Most people who do use them locally just do a prompt every now and then.
This is why I'd like to see a lot more focus on batched inference with lower-end hardware. If you just do a tiny amount of tok/day and can wait for the answer to be computed overnight or so, you don't really need top-of-the-line hardware even for SOTA results.
Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#457I asked it to give me instruction on how to create SSH key and it tried to do it instead of just answering.
Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#458Earlier quoted context omitted.
This is just one model in the Qwen 3.6 series. They will most likely release the other small sizes (not much sense in keeping them proprietary) and perhaps their 122A10B size also, but the flagship 397A17B size seems to have been excluded.
How many people/hackernews can run a 397b param model at home? Probably like 20-30.
Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#459Earlier quoted context omitted.
This is just one model in the Qwen 3.6 series. They will most likely release the other small sizes (not much sense in keeping them proprietary) and perhaps their 122A10B size also, but the flagship 397A17B size seems to have been excluded.
How many people/hackernews can run a 397b param model at home? Probably like 20-30.
Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#460Earlier quoted context omitted.
How many people/hackernews can run a 397b param model at home? Probably like 20-30.
I can (barely, but sustainably) run Q3.5 397B on my Mac Studio with 256GB unified. It cost $10,000 but that's well within reach for most people who are here, I expect.