Earlier quoted context omitted.
For fast inference, you’d be hard pressed to beat an Nvidia RTX 5090 GPU. Check out the HP Omen 45L Max: https://www.hp.com/us-en/shop/pdp/omen-max-45l-gaming-dt-gt2...
I never would have guessed that in 2026, data centers would be measured in Watts and desktop PCs measured in liters.
Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers
91–100 of 286 posts
Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers
#92Smells like hyperbole. A lot of people making such claims don’t seem to have continued real world experience with these models or seem to have very weird standards for what they consider usable. Up until relatively recently, while people had already long been making these claims, it came with the asterisks of „oh, but you can’t practically use more than a few K tokens of context“.
"Create a single page web app scientific RPN calculator" Qwen 3.5 122b/a10b (at q3 using unsloth's dynamic quant) is so far the first model I've tried locally that gets a really usable RPN calculator app. Other models (even larger ones that I can run on my Strix Halo box) tend to either not implement the stack right, have non-functional operation buttons, or most commonly the keypad looks like a Picasso painting (i.e…
Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers
#93Earlier quoted context omitted.
I think the 27B dense model at full precision and 122B MoE at 4- or 6-bit quantization are legitimate killer apps for the 96 GB RTX 6000 Pro Blackwell, if the budget supports it. I imagine any 24 GB card can run the lower quants at a reasonable rate, though, and those are still very good models. Big fan of Qwen 3.5. It actually delivers on some of the hype that the previous wave of open models never lived up to.
I've had good experience with GLM-4.7 and GLM-5.0. How would you compare them with Qwen 3.5? (If you have any experience with them.)
Unsloth's GLM-4.7-Flash-BF16.gguf is quite fast on the 6000, at around 100 t/s, but definitely not as smart as the Qwen 3.5 MoE or dense models of similar size. As far as I'm concerned Qwen 3.5 renders most other open models short of perhaps Kimi 2.5 obsolete for general queries, although other models are still said to be better for local agentic use. That, I haven't tried.
Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers
#94Earlier quoted context omitted.
Well Opus and Gemini are probably running on multiple H200 equivalents, maybe multiple hundreds of thousands of dollars of inference equipment. Local models are inherently inferior; even the best Mac that money can buy will never hold a candle to latest generation Nvidia inference hardware, and the local models, even the largest, are still not quite at the frontier. The ones you can plausibly run on a laptop (where "…
> Well Opus and Gemini are probably running on multiple H200 equivalents, maybe multiple hundreds of thousands of dollars of inference equipment. But if you've got that kind of equipment, you aren't using it to support a single user. It gets the best utilization by running very large batches with massive parallelism across GPUs, so you're going to do that. There is such a thing as a useful middle ground. that may not…
Local models are more than a useful middle ground they are essential and will never go away, I was just addressing the OPs question about why he observed the difference he did. One is an API call to the worlds most advanced compute infrastructure and another is running on a $500 CPU.
Lots of uses for small, medium, and larger models they all have important places!!
Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers
#95Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers
#96I periodically try to run these models on my MBP M3 Max 128G (which I bought with a mind to run local AI). I have a certain deep research question (in a field that is deeply familiar to me) that I ask when I want to gauge model's knowledge. So far Opus 4.6 and Gemini Pro are very satisfactory, producing great answers fairly fast. Gemini is very fast at 30-50 sec, Opus is very detailed and comes at about 2-3 minutes.…
Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers
#97I periodically try to run these models on my MBP M3 Max 128G (which I bought with a mind to run local AI). I have a certain deep research question (in a field that is deeply familiar to me) that I ask when I want to gauge model's knowledge. So far Opus 4.6 and Gemini Pro are very satisfactory, producing great answers fairly fast. Gemini is very fast at 30-50 sec, Opus is very detailed and comes at about 2-3 minutes.…
Also, performance on research-y questions isn't always a good indicator of how the model will do for code generation or agent orchestration.
Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers
#98Earlier quoted context omitted.
Looks at the headline: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers
Yes and Devstral 2 24b q4 is supposed to be 90% as good but it can't even reliably write to a file on my machine. There are the benchmarks, the promises, and what everybody can try at home
Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers
#99Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers
#100I periodically try to run these models on my MBP M3 Max 128G (which I bought with a mind to run local AI). I have a certain deep research question (in a field that is deeply familiar to me) that I ask when I want to gauge model's knowledge. So far Opus 4.6 and Gemini Pro are very satisfactory, producing great answers fairly fast. Gemini is very fast at 30-50 sec, Opus is very detailed and comes at about 2-3 minutes.…
Running local AI models on a laptop is a weird choice. The Mini and especially the Studio form factor will have better cooling, lower prices for comparable specs and a much higher ceiling in performance and memory capacity.