Live data from Hacker News

GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

z.ai

241–250 of 540 posts

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#241

It's live on openrouter now. In my personal benchmark it's bad. So far the benchmark has been a really good indicator of instruction following and agentic behaviour in general. To those who are curious, the benchmark is just the ability of model to follow a custom tool calling format. I ask it to using coding tasks using chat.md [1] + mcps. And so far it's just not able to follow it at all. [1] https://github.com/rus…

Could also be the provider that is bad. Happens way too often on OpenRouter.

I had added z-ai in allow list explicitly and verified that it's the one being used.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#242
post #233

The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…

They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…

It's unclear where the car is currently from your phrasing. If you add that the car is in your garage, it says you'll need to drive to get the car into the wash.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#243
post #233

The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…

They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…

Gemini 3 Pro:

This is a classic logistical puzzle!

Unless you have a very unique way of carrying your vehicle, you should definitely drive.

If you walk there, you'll arrive at the car wash, but your car will still be dirty back at your house. You need to take the car with you to get it washed.

Would you like me to check the weather forecast for $mytown to see if it's a good day for a car wash?

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#244
post #233

The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…

They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…

I just ran this with Gemini 3 Pro, Opus 4.6, and Grok 4 (the models I personally find the smartest for my work). All three answered correctly.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#245

Earlier quoted context omitted.

> Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly some benchmaxxing going on. Agreed. I think the problem is that while they can innovate at algorithms and training efficiency, the human part of RLHF just doesn't scale and they can't afford the massive amount of custom data created a…

Can't they just run the output through a compiler to get feedback? Syntax errors seem easier to get right.

They do. Pretty much all agentic models call linting, compiling and testing tools as part of their flow.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#246
What is truly amazing here is the fact that they trained this entirely on Huawei Ascend chips per reporting [1]. Hence we can conclude the semiconductor to model Chinese tech stack is only 3 months behind the US, considering Opus 4.5 released in November. (Excluding the lithography equipment here, as SMIC still uses older ASML DUV machines) This is huge especially since just a few months ago it was reported that Deepseek were not using Huawei chips due to technical issues [2].

US attempts to contain Chinese AI tech totally failed. Not only that, they cost Nvidia possibly trillions of dollars of exports over the next decade, as the Chinese govt called the American bluff and now actively disallow imports of Nvidia chips as a direct result of past sanctions [3]. At a time when Trump admin is trying to do whatever it can to reduce the US trade imbalance with China.

[1] https://tech.yahoo.com/ai/articles/chinas-ai-startup-zhipu-r...

[2] https://www.techradar.com/pro/chaos-at-deepseek-as-r2-launch...

[3] https://www.reuters.com/world/china/chinas-customs-agents-to...

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#247

Earlier quoted context omitted.

Tim Dettmers had an interesting take on this [1]. Fundamentally, the philosophy is different. >China’s philosophy is different. They believe model capabilities do not matter as much as application. What matters is how you use AI. https://timdettmers.com/2025/12/10/why-agi-will-not-happen/

Sorry, but that's an exceptionally unimpressive article. The crux of his thesis is: >The main flaw is that this idea treats intelligence as purely abstract and not grounded in physical reality. To improve any system, you need resources. And even if a superintelligence uses these resources more effectively than humans to improve itself, it is still bound by the scaling of improvements I mentioned before — linear impro…

Was more mentioning the article about the economic aspect of China vs US in terms of AI.

While I do understand your sentiment, it might be worth noting the author is the author of bitandbytes. Which is one of the first library with quantization methods built in and was(?) one of the most used inference engines. I’m pretty sure transformers from HF still uses this as the Python to CUDA framework

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#248
post #136
post #89

Earlier quoted context omitted.

You're surprised that chinese model makers try to follow chinese law?

This is a classic test to see if the model is censored, as censorship is rarely limited to just one event, which begs the question: what else is censored or outright changed intentionally?

Testing whether a Chinese deep learning model is censored is like testing if water is wet.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#249
post #233

The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…

They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…

How is this riddle relevant to a coding model?
Post reply on HN