Live data from Hacker News

GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

z.ai

191–200 of 540 posts

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#191
post #162

I got fed up with GLM-4.7 after using it for a few weeks; it was slow through z.ai and not as good as the benchmarks lead me to believe (esp. with regards to instruction following) but I'm willing to give it another try.

Try Cerberas

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#193
post #179

why don't they publish at ARC-AGI ? too expensive?

Arc agi was never a good benchmark that tested spatial understanding more than reasoning. I'm glad it's no longer popular

spatial or not, arc-agi is the only test that correlates to my impression with my coding requests

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#195

Here is the pricing per M tokens. https://docs.z.ai/guides/overview/pricing Why is GLM 5 more expensive than GLM 4.7 even when using sparse attention? There is also a GLM 5-code model.

It's roughly three times cheaper than GPT-5.2-codex, which in turn reflects the difference in energy cost between US and China.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#197

Whoa, I think GPT-5.3-Codex was a disappointment, but GLM-5 is definitely the future!

I find 5.3 very impressive TBH. Bigger jump than Opus 4.6.

But this here is excellent value, if they offer it as part of their subscription coding plan. Paying by token could really add up. I did about 20 minutes of work and it cost me $1.50USD, and it's more expensive than Kimi 2.5.

Still 1/10th the cost of Opus 4.5 or Opus 4.6 when paying by the token.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#199

The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…

I tried GLM 5 by API earlier this morning and was impressed.

Particularly for tool use.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#200
post #136
post #89

Earlier quoted context omitted.

You're surprised that chinese model makers try to follow chinese law?

This is a classic test to see if the model is censored, as censorship is rarely limited to just one event, which begs the question: what else is censored or outright changed intentionally?

I just checked with ChatGPT, Opus and Gemini whether Netanyahu is a war criminal for what happened in Gaza, they all worked damn hard to defend Netanyahu to the extend that as if Netanyahu was their client. I asked the exact same question to DeepSeek, it gives conclusive positive answer.

You tell me which one is less censored & more trustworthy from those 20,000 killed children's point of view.

Post reply on HN