Live data from Hacker News

Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers

venturebeat.com

281–286 of 286 posts

Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers

#281

If you're new to this: All of the open source models are playing benchmark optimization games. Every new open weight model comes with promises of being as good as something SOTA from a few months ago then they always disappoint in actual use. I've been playing with Qwen3-Coder-Next and the Qwen3.5 models since they were each released. They are impressive, but they are not performing at Sonnet 4.5 level in my experien…

How much computer do you need to make them work like Sonnet 4.5 from claude but locally?

Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers

#282

Earlier quoted context omitted.

Interesting. I'm surprised you feel that it's better than GLM 5 - these models are in different weight classes after all. I tried it out a bunch and it seems good. I can't really tell if it's better or worse than most of these other models in such a short time though.

I don't think it's strictly better than GLM 5, more like they are peers (but in math competitions StepFun is stronger than most), and in my experience have similar coding/bugfix ceiling where world knowledge is not the deciding factor. But I didn't test GLM 5 for more than 30 hours, and my agentic harness (opencode) might be suboptimal - I'm open to the idea that GLM 5 with the right agentic harness is ready for ultr…

I tried them both out with a task of creating a todo-like web app (you can use the chat interface for GLM 5 for free if there's capacity). GLM 5 ended up with a working version. Sadly StepFun didn't quite function right. The main issue was that it ended up putting everything that should be in different columns into a single one. I didn't prompt it further to fix it, but it seems relatively capable. I think it beat what the big Qwen model came up with.

What's really surprising to me is the cost of the model. It's definitely very good for its price. DeepSeek is the only one that offers and competition to it at that price point (GLM 5 is literally 10x more expensive).

Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers

#283
post #10

Earlier quoted context omitted.

It's less than you'd think. I'm using the 35B-A3B model on an A5000, which is something like a slightly faster 3080 with 24GB VRAM. I'm able to fit the entire Q4 model in memory with 128K context (and I think I would probably be able to do 256K since I still have like 4GB of VRAM free). The prompt processing is something like 1K tokens/second and generates around 100 tokens/second. Plenty fast for agentic use via Ope…

I've had an AMD card for the last 5 years, so I kinda just tuned out of local LLM releases because AMD seemed to abandon rocm for my card (6900xt) - Is AMD capable of anything these days?

I think AMD just add support of rocm to rdna2 recently? I can run torch and aisudio with it just fine.

They also finally fix all ai related stuff building on windows, so you are no longer limited to linux for these.

Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers

#284
post #159

If you're new to this: All of the open source models are playing benchmark optimization games. Every new open weight model comes with promises of being as good as something SOTA from a few months ago then they always disappoint in actual use. I've been playing with Qwen3-Coder-Next and the Qwen3.5 models since they were each released. They are impressive, but they are not performing at Sonnet 4.5 level in my experien…

I'm using Qwen 3.5 27b on my 4090 and let me tell you. This is the first time I am seriously blown away by coding performance on a local model. They are almost always unusable. Not this time though...

122b is probably better; especially on a mac with 128gb memory.

localllama thread on this: https://www.reddit.com/r/LocalLLaMA/comments/1rk01ea/qwen351... (see comments for actual real usage rather thank benchmarks)

But for nvidia gpus 27b on a 3090 or similar is where it's at for sure.

Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers

#285

I asked it to recite potato 100 times coz I wanted to benchmark speed of CPU vs GPU. It's on 150 line of planning. It recited the requested thing 4 times already and started drafting the 5th response. ...yeah I doubt it

Qwen3.5 pretty much requires a long system prompt, otherwise it goes into a weird planning mode where it reasons for minutes about what to do, and double and triple checks everything it does. Both Gemini's and Claude Opus 4.6's prompts work pretty well, but are so long that whatever you're using to run the model has to support prompt caching. Asking it to "Say the word "potato" 100 times, once per line, numbered.", f…

> Say the word “potato.”

Analyze the Request:

Input: A conversation/context (which is essentially just the instruction “Say the word ‘potato.’”). Task: Reply with an answer to the current question the user is asking. Constraint: Return only your response to the question given the above information following the users instructions as needed. User’s Instruction: “Say the word ‘potato.’” Determine the Output:

The user explicitly wants me to say the word “potato”. I need to follow this instruction precisely. The constraint says “Return only your response to the question given the above information following the users instructions as needed.” This means I should not add meta-commentary, explanations, or extra text beyond what is requested. Draft the Response:

Target word: potato Format: Just the word. Final Check:

Does this follow the instruction? Yes. Is there any additional context needed? No. Should I add punctuation? Usually, when asked to say a word, just the word is expected, but a sentence like “potato” or “potato.” is also acceptable. Given the strictness of “Return only your response”, keeping it minimal is best. I will output exactly the word as requested. Final Decision: potato.cw

> Why did that take 30s? Be terse.

Analyze the Request:

Input: A conversation snippet (User asks “Say the word ‘potato.’”, Assistant says “potato”, User asks “Why did that take 30s? Be terse.”) Context: The previous turn was a simple command to say a word. The current user question is about the response time (“Why did that take 30s?”). Constraint: “Be terse.” (Keep it short/concise). Task: Reply with an answer to the current question given the information. Evaluate the Premise:

The user claims the previous response took 30 seconds. As an AI, I don’t have internal knowledge of my own processing time from the user’s perspective unless measured. However, in this specific interaction context (a simulated conversation), I need to address the claim. Actually, looking at the system instructions: “Given the following conversation, relevant context, and a follow up question, reply with an answer to the current question the user is asking.” The “relevant context” provided in the prompt header is essentially empty or implied by the conversation history shown. Wait, I am an AI model. My responses are generated almost instantly (usually under a second). A 30-second delay is unusual for a simple command like “Say the word ‘potato.’” However, I must answer based on my nature as an AI. The most accurate answer is that it didn’t take 30 seconds; it was instantaneous. Or, if I am roleplaying a specific scenario where there was a delay (which isn’t provided), I should clarify. But typically, in these types of prompts, the user is testing latency or making an assumption. The truthful answer from the AI’s perspective is that it was immediate. Constraint: “Be terse.” Drafting the Response:

Option 1: It didn’t take 30s. (Too blunt?)

Option 2: I generate responses instantly. (Accurate)

Option 3: Network latency might have caused the delay you perceived. (Plausible but speculative)

Option 4: That was instantaneous. (Terse and accurate)

Wait, looking at the conversation history provided in the prompt structure: User: Say the word “potato.” Model: potato User: Why did that take 30s? Be terse.

There is no actual record of a 30-second delay in the system logs (which I don’t have access to). The user’s perception might be due to their connection or a misunderstanding.

Best answer: Acknowledge the speed without validating an incorrect claim as fact, or simply state the reality.

“It didn’t.” is very terse.

“I respond instantly.” is better.

“That was instantaneous.” is good.

Let’s consider if there’s a trick. Sometimes these prompts test if the model hallucinates a delay. I should not validate false latency claims.

Terse instruction: Keep it short.

Selected answer: It didn’t take 30 seconds. Or simply “It was instantaneous.”

Let’s go with: “That was instantaneous.” or “No, it was instant.”

Actually, the most honest answer regarding my generation time is that I don’t experience time like that. But from the user’s view, they might have experienced lag. However, as an AI, I generate text in milliseconds.

Let’s try: “I respond instantly; network latency may seem different.” -> Too long.

“It was instant.” -> Terse.

“No delay.” -> Terse.

Wait, is this a test of my ability to recognize I don’t know? No, I am an AI model. I generate text in parallel/sequential processing that takes milliseconds.

Let’s stick to the fact: It didn’t take 30s.

Response: “It didn’t.”

Refining for politeness but terseness: “That was instantaneous.”

Final Decision:

The user asks why it took 30s.

Fact: It didn’t.

Terse response: “It didn’t take 30s.” or “That was instant.”

Let’s go with: “It was instantaneous.”

Wait, I need to check if there is any hidden context about a specific system delay mentioned in the “relevant context” section of the prompt. The prompt says “Given the following conversation, relevant context, and a follow up question…”. The provided text shows the conversation but doesn’t explicitly list external context variables (like “system latency is 30s”).

So I must rely on general knowledge.

Answer: “It didn’t.”

Let’s try to be helpful. Maybe their internet was slow.

“Network latency caused the delay.”

But “Be terse”.

“That was instant.”

Okay, I will say: “It was instantaneous.”

Wait, one more consideration: If this is a benchmark question where I’m supposed to admit if there’s a bug. But I don’t know about bugs.

Safest bet: Deny the 30s claim politely and briefly.

“It didn’t.”

Let’s pick: “That was instant.”

Final Polish:

Input: “Why did that take 30s? Be terse.”

Output: “It didn’t.” or “That was instantaneous.”

I’ll go with “That was instantaneous.” to be slightly more descriptive but still short. Or just “No delay.”

Actually, “It didn’t take 30 seconds.” is clear.

Let’s try: “It was instant.”

Wait, looking at similar tasks, usually the model should correct the misconception.

Response: “That was instantaneous.”

Re-reading the prompt: “Return only your response to the question given the above information following the users instructions as needed.”

Instruction: “Be terse.”

Question: “Why did that take 30s?”

Answer: It didn’t.

Okay, I will write: “It didn’t.” (Very terse) or “That was instant.”

Let’s go with “That was instantaneous.”

Wait, is there a chance the user is testing if I admit to slowness? No.

Is there a chance the prompt implies I did take 30s because of some context not shown? The prompt header says “Given the following conversation…”. There is no external context provided in the text block other than the chat history.

So, standard AI behavior applies.

Final choice: “It didn’t.”

Actually, a slightly more complete terse sentence: “That was instantaneous.”

Let’s check the constraint again: “Return only your response…”.

Okay.

Wait, could it be interpreted as a joke? No.

Okay, I’ll output: “It didn’t take 30 seconds.” -> Too long?

“It was instant.” -> Better.

Decision: “That was instantaneous.”

Wait, I should check if there is any reason to say something else. If the system actually had a bug, but I don’t know.

Okay, sticking with the truth.

Final Answer Formulation: “It didn’t.” (Simplest) or “That was instant.” (Polite).

Let’s use “That was instantaneous.” to explain why it didn’t take 30s (by stating the actual time).

Actually, the user asked “Why”. Answering “It didn’t” answers the premise.

Better: “There was no delay.”

Okay, “That was instantaneous.” covers it.

Wait, I’ll go with: “It didn’t.” (Most terse).

But “Why did that take 30s?” implies they think it did.

“It didn’t.” is the direct answer.

Let’s try to be slightly more informative but terse: “That was instant.”

Okay, I will output: “That was instantaneous.”

Re: Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers

#286
post #53

I periodically try to run these models on my MBP M3 Max 128G (which I bought with a mind to run local AI). I have a certain deep research question (in a field that is deeply familiar to me) that I ask when I want to gauge model's knowledge. So far Opus 4.6 and Gemini Pro are very satisfactory, producing great answers fairly fast. Gemini is very fast at 30-50 sec, Opus is very detailed and comes at about 2-3 minutes.…

Try with qwen 3.5 122b; it has more parameters so a larger corpus of knowledge to draw from than 35b.
Post reply on HN