Live data from Hacker News

GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

z.ai

391–400 of 540 posts

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#391
post #380

What is truly amazing here is the fact that they trained this entirely on Huawei Ascend chips per reporting [1]. Hence we can conclude the semiconductor to model Chinese tech stack is only 3 months behind the US, considering Opus 4.5 released in November. (Excluding the lithography equipment here, as SMIC still uses older ASML DUV machines) This is huge especially since just a few months ago it was reported that Deep…

Where did you read that it was trained on Ascends? I've only seen information suggesting that you can run inference with Ascends, which is obviously a very different thing. The source you link also just says: "The latest model was developed using domestically manufactured chips for inference, including Huawei's flagship Ascend chip and products from leading industry players such as Moore Threads, Cambricon and Kunlun…

I took the "for inference" bit from that sentence you quoted as a qualifier applied to the chips, as in the chips were originally developed for inference but were now used for training too.

Note that Z.ai also publically announced that they trained another model, GLM-Image, entirely on Huawei Ascend silicon a month ago [1].

[1] https://www.scmp.com/tech/tech-war/article/3339869/zhipu-ai-...

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#392
post #386

Earlier quoted context omitted.

https://tech.yahoo.com/ai/articles/chinas-ai-startup-zhipu-r... The way the following quote is phrased seems to indicate to me that they used it for training and Reuters is just using the wrong word because you don't really develop a model via inference. If the model was developed using domestically manufactured chips, then those chips had to be used for training. "The latest model was developed using domestically ma…

Thanks. I'm like 95% sure that you're wrong (as is the parent), and that GLM-5 was trained on NVIDIA GPUs, or at least not on Huawei Ascends. I think so for a few reasons: 1. The Reuters article does explicitly say the model is compatible with domestic chips for inference, without mentioning training. I agree that the Reuters passage is a bit confusing, but I think they mean it was developed to be compatible with Asc…

Fair enough, that makes sense! (2) and (3) especially were convincing to me.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#393
post #386

Earlier quoted context omitted.

Thanks. I'm like 95% sure that you're wrong (as is the parent), and that GLM-5 was trained on NVIDIA GPUs, or at least not on Huawei Ascends. I think so for a few reasons: 1. The Reuters article does explicitly say the model is compatible with domestic chips for inference, without mentioning training. I agree that the Reuters passage is a bit confusing, but I think they mean it was developed to be compatible with Asc…

Fair enough, that makes sense! (2) and (3) especially were convincing to me.

Kudos for changing your mind

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#394

Earlier quoted context omitted.

Custom tool calling formats are iffy in my experience. The models are all reinforcement learned to follow specific ones, so it’s always a battle and feels to me like using the tool wrong. Have you had good results with the other frontier models?

Not the parent commenter, but in my testing, all recent Claudes (4.5 onward) and the Gemini 3 series have been pretty much flawless in custom tool call formats.

Thanks.

I’ve tested local models from Qwen, GLM, and Devstral families.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#395
post #380

Earlier quoted context omitted.

Where did you read that it was trained on Ascends? I've only seen information suggesting that you can run inference with Ascends, which is obviously a very different thing. The source you link also just says: "The latest model was developed using domestically manufactured chips for inference, including Huawei's flagship Ascend chip and products from leading industry players such as Moore Threads, Cambricon and Kunlun…

I took the "for inference" bit from that sentence you quoted as a qualifier applied to the chips, as in the chips were originally developed for inference but were now used for training too. Note that Z.ai also publically announced that they trained another model, GLM-Image, entirely on Huawei Ascend silicon a month ago [1]. [1] https://www.scmp.com/tech/tech-war/article/3339869/zhipu-ai-...

Thanks. I'm like 95% sure that you're wrong, and that GLM-5 was trained on NVIDIA GPUs, or at least not on Huawei Ascends.

As I wrote in another comment, I think so for a few reasons:

1. The z.ai blog post says GML-5 is compatible with Ascends for inference, without mentioning training -- it says they support "deploying GLM-5 on non-NVIDIA chips, including Huawei Ascend, Moore Threads, Cambricon, Kunlun Chip, MetaX, Enflame, and Hygon" -- many different domestic chips. Note "deploying". https://z.ai/blog/glm-5

2. The SCMP piece you linked just says: "Huawei’s Ascend chips have proven effective at training smaller models like Zhipu’s GLM-Image, but their efficacy for training the company’s flagship series of large language models, such as the next-generation GLM-5, was still to be determined, according to a person familiar with the matter."

3. You're right that z.ai trained a small image model on Ascends. They made a big fuss about it too. If they had trained GLM-5 with Ascends, they likely would've shouted it from the rooftops. https://www.theregister.com/2026/01/15/zhipu_glm_image_huawe...

4. Ascends just aren't that good

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#396

Earlier quoted context omitted.

> The U.S. Court of Appeals for the D.C. Circuit has affirmed a district court ruling that human authorship is a bedrock requirement to register a copyright, and that an artificial intelligence system cannot be deemed the author of a work for copyright purposes > The court’s decision in Thaler v. Perlmutter,1 on March 18, 2025, supports the position adopted by the United States Copyright Office and is the latest chap…

It's a fine line that's been drawn, but this ruling says that AI can't own a copyright itself, not that AI output is inherently ineligible for copyright protection or automatically public domain. A human can still own the output from an LLM.

> A human can still own the output from an LLM.

It specifically highlights human authorship, not ownership

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#398

While GLM-5 seems impressive, this release also included lots of new cool stuff! > GLM-5 can turn text or source materials directly into .docx, .pdf, and .xlsx files—PRDs, lesson plans, exams, spreadsheets, financial reports, run sheets, menus, and more. A new type of model has joined the series, GLM-5-Coder. GLM-5 was trained on Huawei Ascend, last time when DeepSeek tried to use this chip, it flopped and they resor…

Not trained in Ascend that is BS. Hopper GPU cluster. Please remove that.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#400

Earlier quoted context omitted.

zAI, minimax and Kimi have plenty of subscriber usage on their own platforms. They get real data just as well. Less or it maybe but it's there.

I'm going to claim that the majority of those users are optimizing for cost and not correctness and therefore the quality of data collected from those sessions is questionable. If you're working on something of consequence, you're not using those platforms. If you're a tinkerer pinching pennies, sure.

ChatGPT, Gemini and Claude are banned in China. Chinese model providers are getting absolutely massive amounts of very valuable user feedback from users in China.
Post reply on HN