What kind of hardware (preferably non-Apple) can run this model? What about 122B?
Qwen3.6-35B-A3B: Agentic coding power, now open to all
71–80 of 563 posts
Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#72What kind of hardware (preferably non-Apple) can run this model? What about 122B?
You can also run those on smaller cards by configuring the number of layers on the GPU. That should allow you to run the Q4/Q5 version on a 4090, or on older cards.
You could also run it entirely on the CPU/in RAM if you have 32GB (or ideally 64GB) of RAM.
The more you run in RAM the slower the inference.
Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#73Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#74Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#75Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#76What do all the numbers 6-35B-A3B mean?
Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#77Earlier quoted context omitted.
This is just one model in the Qwen 3.6 series. They will most likely release the other small sizes (not much sense in keeping them proprietary) and perhaps their 122A10B size also, but the flagship 397A17B size seems to have been excluded.
Is there any source for these claims?
https://x.com/ChujieZheng/status/2039909917323383036
Likely to drive engagement, but the poll excluded the large model size.
Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#78What do all the numbers 6-35B-A3B mean?
The performance/intelligence is said to be about the same as the geometric mean of the total and active parameter counts. So, this model should be equivalent to a dense model with about 10.25 billion parameters.
Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#79Earlier quoted context omitted.
I think its worth noting that if you are paying for electricity Local LLM is NOT free. In most cases you will find that Haiku is cheaper, faster, and better than anything that will run on your local machine.
Electricity (on continental US) is pretty cheap assuming you already have the hardware: Running at a full load of 1000W for every second of the year, for a model that produces 100 tps at 16 cents per kWh, is $1200 USD. The same amount of tokens would cost at least $3,150 USD on current Claude Haiku 3.5 pricing.
Re: Qwen3.6-35B-A3B: Agentic coding power, now open to all
#80Earlier quoted context omitted.
> Close enough No. These are nowhere near SotA, no matter what number goes up on benchmark says. They are amazing for what they are (runnable on regular PCs), and you can find usecases for them (where privacy >> speed / accuracy) where they perform "good enough", but they are not magic. They have limitations, and you need to adapt your workflows to handle them.
Can you share more about what adaptations you made when using smaller models? I'm just starting my exploration of these small models for coding on my 16GB machine (yeah, puny...) and am running into issues where the solution may very well be to reduce the scope of the problem set so the smaller model can handle it.
If you perform the inference locally, there is a huge space of compromise between the inference speed and the quality of the results.
Most open weights models are available in a variety of sizes. Thus you can choose anywhere from very small models with a little more than 1B parameters to very big models with over 750B parameters.
For a given model, you can choose to evaluate it in its native number size, which is normally BF16, or in a great variety of smaller quantized number sizes, in order to fit the model in less memory or just to reduce the time for accessing the memory.
Therefore, if you choose big models without quantization, you may obtain results very close to SOTA proprietary models.
If you choose models so small and so quantized as to run in the memory of a consumer GPU, then it is normal to get results much worse than with a SOTA model that is run on datacenter hardware.
Choosing to run models that do not fit inside the GPU memory reduces the inference speed a lot, and choosing models that do not fit even inside the CPU memory reduces the inference speed even more.
Nevertheless, slow inference that produces better results may reduce the overall time for completing a project, so one should do a lot of experiments to determine an appropriate compromise.
When you use your own hardware, you do not have to worry about token cost or subscription limits, which may change the optimal strategy for using a coding assistant. Moreover, it is likely that in many cases it may be worthwhile to use multiple open-weights models for the same task, in order to choose the best solution.
For example, when comparing older open-weights models with Mythos, by using appropriate prompts all the bugs that could be found by Mythos could also be found by old models, but the difference was that Mythos found all the bugs alone, while with the free models you had to run several of them in order to find all bugs, because all models had different strengths and weaknesses.
(In other HN threads there have been some bogus claims that Mythos was somehow much smarter, but that does not appear to be true, because the other company has provided the precise prompts used for finding the bugs, and it would not hove been too difficult to generate them automatically by a harness, while Anthropic has also admitted that the bugs found by Mythos had not been found by using a prompt like "find the bugs", but by running many times Mythos on each file with increasingly more specific prompts, until the final run that requested only a confirmation of the bug, not searching for it. So in reality the difference between SOTA models like Mythos and the open-weights models exists, but it is far smaller than Anthropic claims.)