Earlier quoted context omitted.
They are trained to respond to certain topics in a way that does not align with real world evidence. Pretty much the opposite of what you want in such a tool. This is trivial to test and verify yourself. Just pick any topic you think has a chance of being censored. You can do the same on American models and compare results.
I strongly agree, that this is an issue that needs to be adressed, but for everyday coding tasks it wont matter in 99.9 percent of cases.
Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers
221–230 of 286 posts
Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers
#222If you're new to this: All of the open source models are playing benchmark optimization games. Every new open weight model comes with promises of being as good as something SOTA from a few months ago then they always disappoint in actual use. I've been playing with Qwen3-Coder-Next and the Qwen3.5 models since they were each released. They are impressive, but they are not performing at Sonnet 4.5 level in my experien…
Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers
#223I periodically try to run these models on my MBP M3 Max 128G (which I bought with a mind to run local AI). I have a certain deep research question (in a field that is deeply familiar to me) that I ask when I want to gauge model's knowledge. So far Opus 4.6 and Gemini Pro are very satisfactory, producing great answers fairly fast. Gemini is very fast at 30-50 sec, Opus is very detailed and comes at about 2-3 minutes.…
I've seen reports of qwen3.5-35b-a3b spending a ton of time reasoning if the context window is nearly empty-- supposedly it reasons less if you provide a long system prompt or some file contents, like if you use it in a coding agent. I'm too GPU-poor to run it, but r/LocalLLaMa is full of people using it.
On the plus side, it did figure out the question even without the first sentence that's intended as a bit of a giveaway.
Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers
#224If you're new to this: All of the open source models are playing benchmark optimization games. Every new open weight model comes with promises of being as good as something SOTA from a few months ago then they always disappoint in actual use. I've been playing with Qwen3-Coder-Next and the Qwen3.5 models since they were each released. They are impressive, but they are not performing at Sonnet 4.5 level in my experien…
> they always disappoint in actual use. I’ve switched to using Kimi 2.5 for all of my personal usage and am far from disappointed. Aside from being much cheaper than the big names (yes, I’m not running it locally, but like that I could) it just works and isn’t a sycophant. Nice to get coding problems solved without any “That’s a fantastic idea!”/“great point” comments. At least with Kimi my understanding is that beat…
Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers
#225I periodically try to run these models on my MBP M3 Max 128G (which I bought with a mind to run local AI). I have a certain deep research question (in a field that is deeply familiar to me) that I ask when I want to gauge model's knowledge. So far Opus 4.6 and Gemini Pro are very satisfactory, producing great answers fairly fast. Gemini is very fast at 30-50 sec, Opus is very detailed and comes at about 2-3 minutes.…
The reality in ML is that small models can perform better at a narrow problem set than large ones.
The key is the narrow problem set. Opus can write you a poem, create a shopping list, and analyze your massive code base.
We trained our model to only focus on coding with our specific agent harness, tools, and context engine. And it’s small enough to fit on an M2 16GB. It’s as good as sonnet 4.5 and way better than qwen3.5:35b-a3b
Our beta will be out soon / rig.ai
Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers
#226Earlier quoted context omitted.
I pay for copilot to access anthropic, google and openai models. Claude code always give me rate limits. Claude through copilot is a bit slow, but copilot has constant network request issues or something, but at least I don't get rate limited as often. At least local models always work, is faster (50+ tps with qwen3.5 35b a4b on a 4090) and most importantly never hit a rate limit.
> Claude code always give me rate limits > 50+ tps with qwen3.5 35b a4b on a 4090 But qwen3.5 35b is worse than even Claude Haiku 4.5. You could switch your Claude Code to use Haiku and never hit rate limits. Also gets similar 50tps.
My goto proprietary model in copilot for general tasks is gemini 3 flash which is priced the same as haiku.
The qwen model is in my experience close to gemini 3 flash, but gemini flash is still better.
Maybe it's somewhat related to what we're using them for. In my case I'm mostly using llms to code Lua. One case is a typed luajit language and the other is a 3d luajit framework written entirely in luajit.
I forgot exactly how many tps i get with qwen, but with glm 4.7 flash which is really good (to be local) gets me 120tps and a 120k context.
Don't get me wrong, proprietary models are superior, but local models are getting really good AND useful for a lot of real work.
Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers
#227Earlier quoted context omitted.
The cheapest option is two 3060 12G cards. You'll be able to fit the Q4 of the 27B or 35B with an okay context window. If you want to spend twice as much for more speed, get a 3090/4090/5090. If you want long context, get two of them. If you have enough spare cash to buy a car, get an RTX Ada with 96G VRAM.
Rtx 6000 pro Blackwell, not ada, for 96GB.
The names are so good and not repetitious.
No not the RTX 6000. No not the A6000...
Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers
#228If you're new to this: All of the open source models are playing benchmark optimization games. Every new open weight model comes with promises of being as good as something SOTA from a few months ago then they always disappoint in actual use. I've been playing with Qwen3-Coder-Next and the Qwen3.5 models since they were each released. They are impressive, but they are not performing at Sonnet 4.5 level in my experien…
"When a measure becomes a target, it ceases to be a good measure." Goodhart's law shows up with people, in system design, in processor design, in education... Models are going to be over-fit to the tests unless scruples or practical application realities intervene. It's a tale as old as machine learning.
But there's a problem with that: of course the existence of the statistical measure itself is very much a link between all those individual facts. In other words: if there is ANY causal link between the statistical measure and the events measured ... it has now become bullshit (because the law of large numbers doesn't apply anymore).
So let's put it in practice, say there's a running contest, and you display the minimum, maximum and average time of all runners that have had their turns. We all know what happens: of course the result is that the average trends up. And yet, that's exactly what statistics guarantees won't happen. The average should go up and down with roughly 50% odds when a new runner is added. This is because showing the average causes behavior changes in the next runner.
This means, of course, that basing a decision on something as trivial as what the average running time was last year can only be mathematically defensible ONCE. The second time the average is wrong, and you're basing your decision on wrong information.
But of course, not only will most people actually deny this is the case, this is also how 99.9% of human policy making works. And it's mathematically wrong! Simple, fast ... and wrong.
Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers
#229Earlier quoted context omitted.
I strongly agree, that this is an issue that needs to be adressed, but for everyday coding tasks it wont matter in 99.9 percent of cases.
I found that as well. It replaces American propaganda with Chinese propaganda.
Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers
#230If you're new to this: All of the open source models are playing benchmark optimization games. Every new open weight model comes with promises of being as good as something SOTA from a few months ago then they always disappoint in actual use. I've been playing with Qwen3-Coder-Next and the Qwen3.5 models since they were each released. They are impressive, but they are not performing at Sonnet 4.5 level in my experien…
I'm using Qwen 3.5 27b on my 4090 and let me tell you. This is the first time I am seriously blown away by coding performance on a local model. They are almost always unusable. Not this time though...