Live data from Hacker News

GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

z.ai

151–160 of 540 posts

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#151
post #29

Grey market fast-follow via distillation seems like an inevitable feature of the near to medium future. I've previously doubted that the N-1 or N-2 open weight models will ever be attractive to end users, especially power users. But it now seems that user preferences will be yet another saturated benchmark, that even the N-2 models will fully satisfy. Heck, even my own preferences may be getting saturated already. Op…

Just to say - 4.6 really shines on working longer without input. It feels to me like it gets twice as far. I would not want to go back.

If that's what they're tuning for, that's just not what I want. So I'm glad I switched off of Anthropic.

What teams of programmers need, when AI tooling is thrown into the mix, is more interaction with the codebase, not less. To build reliable systems the humans involved need to know what was built and how.

I'm not looking for full automation, I'm looking for intelligence and augmentation, and I'll give my money and my recommendation as team lead / eng manager to whatever product offers that best.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#152
post #5

It's looking like we'll have Chinese OSS to thank for being able to host our own intelligence, free from the whims of proprietary megacorps. I know it doesn't make financial sense to self-host given how cheap OSS inference APIs are now, but it's comforting not being beholden to anyone or requiring a persistent internet connection for on-premise intelligence. Didn't expect to go back to macOS but they're basically the…

I'm not sure being beholden to the whims of the Chinese Communist Party is an iota better than the whims of proprietary megacorps, especially given this probably will become part of a megacorp anyway.

It seems you missed the point entirely once you saw the word "Chinese". The point isn't that the models are from China. It's that the weights are open. You can download the weights and finetune them yourself. Nobody is beholden to anything.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#153

Earlier quoted context omitted.

As I promised earlier: https://news.ycombinator.com/item?id=46781777 "I will save this for the future, when people complain about Chinese open models and tell me: But this Chinese LLM doesn't respond to question about Tianmen square." Please stop using Tianmen question as an example to evaluate the company or their models: https://news.ycombinator.com/item?id=46779809

That's just whataboutism. Why shouldn't people talk about the various ideological stances embedded in different LLMs?

Why do we hear censorship concerns only when it comes Chinese models? Why don't we hear similar stances when Claude or OpenAI releases models?

We either set the bar and judge both, or don't complain about censorship

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#154
post #5

It's looking like we'll have Chinese OSS to thank for being able to host our own intelligence, free from the whims of proprietary megacorps. I know it doesn't make financial sense to self-host given how cheap OSS inference APIs are now, but it's comforting not being beholden to anyone or requiring a persistent internet connection for on-premise intelligence. Didn't expect to go back to macOS but they're basically the…

[deleted]

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#155
post #115

Bought some API credits and ran it through opencode (model was "GLM 5"). Pretty impressed, it did good work. Good reasoning skills and tool use. Even in "unfamiliar" programming languages: I had it connect to my running MOO and refactor and rewrite some MOO (dynamic typed OO scripting language) verbs by MCP. It made basically no mistakes with the programming language despite it being my own bespoke language & runtime…

when i look at the prices these people are offering, and also the likes of kimi, and I wonder how are openAI, anthropic and google going to justify billions of dollars of investment? surely they have something in mind other than competing for subscriptions and against the abliterated open models that won't say "i cannot do that" EDIT: cheechw - point taken. I'm very sceptical of that business model also, as it's fair…

[dead]

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#156

Earlier quoted context omitted.

As I promised earlier: https://news.ycombinator.com/item?id=46781777 "I will save this for the future, when people complain about Chinese open models and tell me: But this Chinese LLM doesn't respond to question about Tianmen square." Please stop using Tianmen question as an example to evaluate the company or their models: https://news.ycombinator.com/item?id=46779809

Neither should be censoring objective reality. Why defend it on either side?

> Neither should be censoring objective reality.

100% agree!

But Chinese model releases are treated unfairly all the time when they release new model, as if Tianmen response indicates that we can use the model for coding tasks.

We should understand their situation and don't judge for obvious political issue. Its easy to judge people working hard over there, because they are confirming to the political situation and don't want to kill their company.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#157
post #91

Earlier quoted context omitted.

> doesn't make financial sense to self-host I guess that's debatable. I regularly run out of quota on my claude max subscription. When that happens, I can sort of kind of get by with my modest setup (2x RTX3090) and quantized Qwen3. And this does not even account for privacy and availability. I'm in Canada, and as the US is slowly consumed by its spiral of self-destruction, I fully expect at some point a digital iron…

Your $5,000 PC with 2 GPUs could have bought you 2 years of Claude Max, a model much more powerful and with longer context. In 2 years you could make that investment back in pay raise.

> In 2 years you could make that investment back in pay raise.

Could you elaborate? I fail to grasp the implication here.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#159
post #149
post #136

Earlier quoted context omitted.

This is a classic test to see if the model is censored, as censorship is rarely limited to just one event, which begs the question: what else is censored or outright changed intentionally?

> which begs the question: what else is censored or outright changed intentionally? So like every other frontier model that has post training to add safeguards in accordance with local norms. Claude won't help you hotwire a car. Gemini won't write you erotic novels. GPT won't talk about suicide or piracy. etc etc >This is a classic test It's a gotcha question with basic zero real world relevance I'd prefer models to…

The problem with censorship isn't that it degrades performance. The problem is that if the censorship is unilaterally dictated by a government then it becomes a tool for suppression, especially as people use AI more and more for their primary source of information.

A company might choose to avoid erotica because it clashes with their brand, or avoid certain topics because they're worried about causing harms. That is very different than centralized, unilateral control over all information sources.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#160
I'd say that they're super confident about the GLM-5 release, since they're directly comparing it with Opus 4.5 and don't mention Sonnet 4.5 at all.

I am still waiting if they'd launch GLM-5 Air series,which would run on consumer hardware.

Post reply on HN