Live data from Hacker News

GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

z.ai

291–300 of 540 posts

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#291

The amount of times benchmarks of competitors said something is close to Claude and it was remotely close in practice in the past year: 0

I honestly feel like people are brainwashed by anthropic propaganda when it comes to claude, I think codex is just way better and kimi 2.5 (and I think glm 5 now) are perfectly fine for a claude replacement.

So much money is on the line for US super scalers that they probably pay for ‘pushes’ on social media. Maybe Chinese companies are doing the same.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#292

Earlier quoted context omitted.

I think more people should spend time talking about this with American models, yeah. If you're interested in that then maybe that can be you. It doesn't have to be the same exact people talking about everything, that's the nice thing about forums. Find your own topic that American models consistently lie or freeze on that Chinese models don't and post about it.

I don't want to criticise models for things they're not being trained on or constraints companies have. None of the companies said our models don't hallucinate and we always have right facts. For example, * I am not expecting Gemini 3 Flash to cure cancer and constantly criticising them for that * Or I am not expecting Mistral to outcompete OpenAI/Claude on their each release, because talent density and capital is ob…

I think most people feel differently about an emergent failure in a model vs one that's been deliberately engineered in for ideological reasons.

It's not like Chinese models just happen to refuse to talk about the topic, it trips guardrails that have been intentionally placed there, just as much as Claude has guardrails against telling you how to make sarin gas.

eg ChatGPT used to have an issue where it steadfastly refused to make any "political" judgments, which led it to genocide denial or minimization- "could genocide be justifiable" to which sometimes it would refuse to say "no." Maybe it still does this, I haven't checked, but it seemed very clearly a product of being strongly biased against being "political", which is itself an ideology and worth talking about.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#293

Why are we not comparing to opus 4.6 and gpt 5.3 codex... Honestly these companies are so hard to takes seriously with these release details. If it's an open source model and you're only comparing open source - cool. If you're not top in your segment, maybe show how your token cost and output speed more than make up for that. Purposely showing prior-gen models in your release comparison immediately discredits you in…

I feel like you're over reacting.

They're comparing against 5.2 xhigh, which is arguably better than 5.3. The latest from openai isn't smarter, it's slightly dumber, just much faster.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#294

Been using GLM-4.7 for a couple weeks now. Anecdotally, it’s comparable to sonnet, but requires a little bit more instruction and clarity to get things right. For bigger complex changes I still use anthropic’s family, but for very concise and well defined smaller tasks the price of GLM-4.7 is hard to beat.

This aligns very closely with my experience.

When left to its own devices, GLM-4.7 frequently tries to build the world. It's also less capable at figuring out stumbling blocks on its own without spiralling.

For small, well-defined tasks, it's broadly comparable to Sonnet.

Given how incredibly cheap it is, it's useful even as a secondary model.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#295

Earlier quoted context omitted.

> Going from GLM-4.7 to something comparable to 4.5 or 5.2 would be an absolutely crazy improvement. Before you get too excited, GLM-4.7 outperformed Opus 4.5 on some benchmarks too - https://www.cerebras.ai/blog/glm-4-7 See the LiveCodeBench comparison The benchmarks of the open weights models are always more impressive than the performance. Everyone is competing for attention and market share so the incentives to b…

Sure. My sole point is that calling Opus 4.5 and GPT-5.2 "last generation models" is discounting how good they are. In fact, in my experience, Opus 4.6 isn't much of an improvement over 4.5 for agentic coding. I'm not immediately discounting Z.ai's claims because they showed with GLM-4.7 that they can do quite a lot with very little. And Kimi K2.5 is genuinely a great model, so it's possible for Chinese open-weight m…

From a user perspective, I would consider Opus 4.6 somewhat of a regression. You can exhaust your the five hour limit in less than half an hour on, and I used up the weekly limit in just two days. The outputs did not feel significantly better than Opus 4.5 and that only feels smarter than Sonnet by degrees. This is running a single session on a pro plan. I don’t get paid to program, so API cost matter to me. The experience was irritating enough to make me start looking for an alternative, and maybe GLM is the way to go for hobby users.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#296

The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…

What a strangely hostile statement on an open weight model. Running like 20 benchmark evaluations isn't trivial by itself, and even updating visuals and press statements can take a few days at a tech company. It's literally been 5 days since this "new generation" of models released. GPT-5.3(-codex) can't even be called via API, so it's impossible to test for some benchmarks.

I notice the people who endlessly praise closed-source models never actually USE open weight models, or assume their drop-in prompting methods and workflow will just work for other model families. Especially true for SWEs who used Claude Code first and now think every other model is horrible because they're ONLY used to prompting Claude. It's quite scary to see how people develop this level of worship for a proprietary product that is openly distrusting of users. I am not saying this is true or not of the parent poster, but something I notice in general.

As someone who uses GLM-4.7 a good bit, it's easily at Sonnet 4.5 tier - have not tried GLM-5 but it would be surprising if it wasn't at Opus 4.5 level given the massive parameter increase.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#297

Been using GLM-4.7 for a couple weeks now. Anecdotally, it’s comparable to sonnet, but requires a little bit more instruction and clarity to get things right. For bigger complex changes I still use anthropic’s family, but for very concise and well defined smaller tasks the price of GLM-4.7 is hard to beat.

Anecdotal, but I've been locked to Sonnet for the past 6-8 months just because they always seem to introduce throttling bugs with Opus where it starts to devour tokens or falls over. Very interested once open models close the gap to about 6 months.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#298
post #215

Pelican generated via OpenRouter: https://gist.github.com/simonw/cc4ca7815ae82562e89a9fdd99f07... Solid bird, not a great bicycle frame.

How many pelican riding bicycle SVGs were there before this test existed? What if the training data is being polluted with all these wonky results...

You're correct. It's not as useful as it (ever?) was as a measure of performance...but it's fun and brings me joy.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#299
I predict a new speculative market will emerge where adherents buy and sell misween coded companies.

Betting on whether they can actually perform their sold behaviors.

Passing around code repositories for years without ever trying to run them, factory sealed.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#300

It might be impressive on benchmarks, but there's just no way for them to break through the noise from the frontier models. At these prices they're just hemorrhaging money. I can't see a path forward for the smaller companies in this space.

I expect that the reason for their existence is political rather than financial (though I have no idea how that's structured.)

It's a big deal that open-source capability is less than a year behind frontier models.

And I'm very, very glad it is. A world in which LLM technology is exclusive and proprietary to three companies from the same country is not a good world.

Post reply on HN