Live data from Hacker News

Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers

venturebeat.com

91–100 of 286 posts

Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers

#91

Earlier quoted context omitted.

For fast inference, you’d be hard pressed to beat an Nvidia RTX 5090 GPU. Check out the HP Omen 45L Max: https://www.hp.com/us-en/shop/pdp/omen-max-45l-gaming-dt-gt2...

I never would have guessed that in 2026, data centers would be measured in Watts and desktop PCs measured in liters.

The Omen was neigh.

Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers

#92
post #68

Smells like hyperbole. A lot of people making such claims don’t seem to have continued real world experience with these models or seem to have very weird standards for what they consider usable. Up until relatively recently, while people had already long been making these claims, it came with the asterisks of „oh, but you can’t practically use more than a few K tokens of context“.

"Create a single page web app scientific RPN calculator" Qwen 3.5 122b/a10b (at q3 using unsloth's dynamic quant) is so far the first model I've tried locally that gets a really usable RPN calculator app. Other models (even larger ones that I can run on my Strix Halo box) tend to either not implement the stack right, have non-functional operation buttons, or most commonly the keypad looks like a Picasso painting (i.e…

We tend to find Qwen3-Coder-Next better at coding at least on our anecdotal examples from our codebases. It's somewhat better at tool calling, maybe the current templates for Qwen3.5 are still not enjoying as "mature" support as Qwen3 on vllm. I can say in my team MiniMax2.5 is the currently favorite.

Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers

#93

Earlier quoted context omitted.

I think the 27B dense model at full precision and 122B MoE at 4- or 6-bit quantization are legitimate killer apps for the 96 GB RTX 6000 Pro Blackwell, if the budget supports it. I imagine any 24 GB card can run the lower quants at a reasonable rate, though, and those are still very good models. Big fan of Qwen 3.5. It actually delivers on some of the hype that the previous wave of open models never lived up to.

I've had good experience with GLM-4.7 and GLM-5.0. How would you compare them with Qwen 3.5? (If you have any experience with them.)

No experience with 5 and not much with 4.7, but they both have quite a few advocates over on /r/localllama.

Unsloth's GLM-4.7-Flash-BF16.gguf is quite fast on the 6000, at around 100 t/s, but definitely not as smart as the Qwen 3.5 MoE or dense models of similar size. As far as I'm concerned Qwen 3.5 renders most other open models short of perhaps Kimi 2.5 obsolete for general queries, although other models are still said to be better for local agentic use. That, I haven't tried.

Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers

#94

Earlier quoted context omitted.

Well Opus and Gemini are probably running on multiple H200 equivalents, maybe multiple hundreds of thousands of dollars of inference equipment. Local models are inherently inferior; even the best Mac that money can buy will never hold a candle to latest generation Nvidia inference hardware, and the local models, even the largest, are still not quite at the frontier. The ones you can plausibly run on a laptop (where "…

> Well Opus and Gemini are probably running on multiple H200 equivalents, maybe multiple hundreds of thousands of dollars of inference equipment. But if you've got that kind of equipment, you aren't using it to support a single user. It gets the best utilization by running very large batches with massive parallelism across GPUs, so you're going to do that. There is such a thing as a useful middle ground. that may not…

Batching helps with efficiency but you can’t fit opus into anything less than hundreds of thousands of dollars in equipment

Local models are more than a useful middle ground they are essential and will never go away, I was just addressing the OPs question about why he observed the difference he did. One is an API call to the worlds most advanced compute infrastructure and another is running on a $500 CPU.

Lots of uses for small, medium, and larger models they all have important places!!

Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers

#95
SWE chart is missing Claude on front page, interesting way to present your data. Mix and match at will. Grown up people showing public school level sneakiness. That fact alone disqualifies your LL. Business/marketing leaders usually are brighter than average developers... so there.

Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers

#96
post #53

I periodically try to run these models on my MBP M3 Max 128G (which I bought with a mind to run local AI). I have a certain deep research question (in a field that is deeply familiar to me) that I ask when I want to gauge model's knowledge. So far Opus 4.6 and Gemini Pro are very satisfactory, producing great answers fairly fast. Gemini is very fast at 30-50 sec, Opus is very detailed and comes at about 2-3 minutes.…

You're not doing anything wrong. The Chinese models are not as good as advertised. Surprise surprise!

Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers

#97
post #53

I periodically try to run these models on my MBP M3 Max 128G (which I bought with a mind to run local AI). I have a certain deep research question (in a field that is deeply familiar to me) that I ask when I want to gauge model's knowledge. So far Opus 4.6 and Gemini Pro are very satisfactory, producing great answers fairly fast. Gemini is very fast at 30-50 sec, Opus is very detailed and comes at about 2-3 minutes.…

Try the 27B dense model. It will likely do much better than the 35b MoE with only 3B active experts.

Also, performance on research-y questions isn't always a good indicator of how the model will do for code generation or agent orchestration.

Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers

#98
post #89

Earlier quoted context omitted.

Looks at the headline: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers

Yes and Devstral 2 24b q4 is supposed to be 90% as good but it can't even reliably write to a file on my machine. There are the benchmarks, the promises, and what everybody can try at home

maybe a harness problem?

Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers

#100
post #53

I periodically try to run these models on my MBP M3 Max 128G (which I bought with a mind to run local AI). I have a certain deep research question (in a field that is deeply familiar to me) that I ask when I want to gauge model's knowledge. So far Opus 4.6 and Gemini Pro are very satisfactory, producing great answers fairly fast. Gemini is very fast at 30-50 sec, Opus is very detailed and comes at about 2-3 minutes.…

Running local AI models on a laptop is a weird choice. The Mini and especially the Studio form factor will have better cooling, lower prices for comparable specs and a much higher ceiling in performance and memory capacity.

I have a laptop already, so that's what I'm going to use.
Post reply on HN