Live data from Hacker News

GLM 4.5 with Claude Code

docs.z.ai

51–60 of 90 posts

Re: GLM 4.5 with Claude Code

#51
It's weird, but using Claude Code (CC) with GLM 4.5 from Z.ai, I spent $5 just starting CC and asking it to read the spec and guideline files. This on an API that advertises $0.6 per million tokens input, and $2.2 per million tokens output.

Then, I noticed that after the original prompt, any further prompt I gave CC (other that approving its actions) did not have any effect at all on it. Yes, any prompt is totally ignored. I have to stop CC and restart with the new prompt for it to take effect.

This is really strange to me. Does CC have some way of detecting when it is running against an API other than Anthropic's? The massive cost and/or token usage and crippled agentic mode is clearly not due to GLM 4.5 not being capable. Either CC is crippled in some way when using a 3rd party API, or else Z.ai's API is somehow not working properly.

Re: GLM 4.5 with Claude Code

#52

Earlier quoted context omitted.

So what's the deal with Chutes and all the throttling and errors. Seems like users are losing their minds over this.. at least from all the reddit threads I'm seeing

What's chutes?

Cheap provider on OpenRouter:

https://openrouter.ai/provider/chutes

Re: GLM 4.5 with Claude Code

#53
post #40

Earlier quoted context omitted.

you can use any model with Claude code thanks to https://github.com/musistudio/claude-code-router but in my testing other models do not work well, looks like prompts are either very optimized for Claude, or other models are just not great yet with such agentic environment I was especially disappointed with grok code. it is very fast as advertised but in generating spaces and new lines in function calling until it hit…

You don't need claude code router to use GLM, just set the env var to the GLM url. Also, I generally advise people not to bother with claude code router, Bifrost can do the same job and it's much better software.

I wasn't aware that there is an alternative

Quick glance over readme suggest it's only openai compatible but I also found HN post [1] explaining use of claude code with ollama

But anyways, claude-code-router have advantage of allowing request transformers, those are required for getting GitHub copilot as provider and grok-code limitations to messages format.

[1] https://news.ycombinator.com/item?id=45127509

Re: GLM 4.5 with Claude Code

#54
post #4

Okay, I'm going to try it, but why didn't you link the information on how to integrate it with Claude Code: https://docs.z.ai/scenario-example/develop-tools/claude Chinese software always has such a design language: - prepaid and then use credit to subscribe - strange serif font - that slider thing for captcha But I'm going to try it out now.

> prepaid and then use credit to subscribe

This is mainly because Chinese online payment infrastructure didn't have good support for subscriptions or auto payments (at least until relatively recently) so this pattern is the norm

Re: GLM 4.5 with Claude Code

#55
post #4

Okay, I'm going to try it, but why didn't you link the information on how to integrate it with Claude Code: https://docs.z.ai/scenario-example/develop-tools/claude Chinese software always has such a design language: - prepaid and then use credit to subscribe - strange serif font - that slider thing for captcha But I'm going to try it out now.

> prepaid and then use credit to subscribe This is mainly because Chinese online payment infrastructure didn't have good support for subscriptions or auto payments (at least until relatively recently) so this pattern is the norm

It's more of a culture thing. People just hate the concept of "idk how much I'm going to pay let's just try this and find out later".

Also people would be confused as they expect things to be prepaid, so if you let them use the service they'd think it's a free trial or something, unless you literally put very big, clear price tag and require like triple confirmation. If not and you ask them to pay later they would perceive this as unfair deceptive tricks, and may scam you by report the loss of their credit card (!), because apparently disputing transactions in China is super hard.

Re: GLM 4.5 with Claude Code

#56

Not just Claude Code. Their plans $3 and $15 plans work even better with tools like Roo Code. After Claude models have recently become dumb, I switched to Qwen3-Coder (there's a very generous free tier) and GLM4.5, and I'm not looking back.

How do you use qwen3 coder free tier?

Re: GLM 4.5 with Claude Code

#57
post #49

Earlier quoted context omitted.

> I get better results from Qwen 3 coder 30b a3b locally than I get from Qwen 3 Coder 480b through open router. I'm really concerned that some of the providers are using quantized versions of the models so they can run more models per card and larger batches of inference. This doesn't match my experience precisely, but I've definitely had cases where some of the providers had consistently worse output for the same mo…

Interesting. Thanks for sharing. What about qwen3-coder on Cerebras? I'm happy to pay the $50 for the speed as long as results are good. How does it compare with glm-4.5?

I wish that Cerebras had a direct pay per use API option instead of pushing you towards OpenRouter and HuggingFace (the former sometimes throws 429, so either the speed is great, or there is no speed): https://www.cerebras.ai/pricing but I imagine that for most folks their subscription would be more than enough!

As for how Qwen3 Coder performs, there's always SWE-bench: https://www.swebench.com/

By the numbers:

  * it sits between Gemini 2.5 Pro and GPT-5 mini
  * it beats out Kimi K2 and the older Claude Sonnet 3.7
  * but loses out to Claude Sonnet 4 and GPT-5
Personally, I find it sufficient for most tasks (from recommendations and questions to as close to vibe coding as I get) on a technical level. GLM 4.5 isn't on the site at the time of writing this, but they should match one another pretty closely. Feeling wise, I still very much prefer Sonnet 4 to everything else, but it's both expensive and way slower than Cerebras (not even close).

Update: also seems like the Growth plan on their page says "Starting from 1500 USD / month" which is a bit silly when the new cheapest subscription is 50 USD / month.

Re: GLM 4.5 with Claude Code

#58

Not just Claude Code. Their plans $3 and $15 plans work even better with tools like Roo Code. After Claude models have recently become dumb, I switched to Qwen3-Coder (there's a very generous free tier) and GLM4.5, and I'm not looking back.

How do you use qwen3 coder free tier?

https://x.com/Alibaba_Qwen/status/1953835877555151134

Re: GLM 4.5 with Claude Code

#59

Not just Claude Code. Their plans $3 and $15 plans work even better with tools like Roo Code. After Claude models have recently become dumb, I switched to Qwen3-Coder (there's a very generous free tier) and GLM4.5, and I'm not looking back.

How do you use qwen3 coder free tier?

You can install Qwen Code (the Gemini CLI fork) and use OAuth for authentication. That will give you 2000 free requests per day, no token limits.

RooCode can use OAuth as well (but you have to install Qwen Code first).

Re: GLM 4.5 with Claude Code

#60

Anthropic can't compete with this on cost. They're probably bleeding money as it is. But they can sort of compete on model quality, by no longer dumbing down their models. That'll be expensive too, but it's a lever they have.

One of the articles posted here a week ago says Claude Code has about a ~20x margin on inference. So they can compete on cost if they want.
Post reply on HN