The amount of times benchmarks of competitors said something is close to Claude and it was remotely close in practice in the past year: 0
I honestly feel like people are brainwashed by anthropic propaganda when it comes to claude, I think codex is just way better and kimi 2.5 (and I think glm 5 now) are perfectly fine for a claude replacement.
GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
291–300 of 540 posts
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#292Earlier quoted context omitted.
I think more people should spend time talking about this with American models, yeah. If you're interested in that then maybe that can be you. It doesn't have to be the same exact people talking about everything, that's the nice thing about forums. Find your own topic that American models consistently lie or freeze on that Chinese models don't and post about it.
I don't want to criticise models for things they're not being trained on or constraints companies have. None of the companies said our models don't hallucinate and we always have right facts. For example, * I am not expecting Gemini 3 Flash to cure cancer and constantly criticising them for that * Or I am not expecting Mistral to outcompete OpenAI/Claude on their each release, because talent density and capital is ob…
It's not like Chinese models just happen to refuse to talk about the topic, it trips guardrails that have been intentionally placed there, just as much as Claude has guardrails against telling you how to make sarin gas.
eg ChatGPT used to have an issue where it steadfastly refused to make any "political" judgments, which led it to genocide denial or minimization- "could genocide be justifiable" to which sometimes it would refuse to say "no." Maybe it still does this, I haven't checked, but it seemed very clearly a product of being strongly biased against being "political", which is itself an ideology and worth talking about.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#293Why are we not comparing to opus 4.6 and gpt 5.3 codex... Honestly these companies are so hard to takes seriously with these release details. If it's an open source model and you're only comparing open source - cool. If you're not top in your segment, maybe show how your token cost and output speed more than make up for that. Purposely showing prior-gen models in your release comparison immediately discredits you in…
They're comparing against 5.2 xhigh, which is arguably better than 5.3. The latest from openai isn't smarter, it's slightly dumber, just much faster.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#294Been using GLM-4.7 for a couple weeks now. Anecdotally, it’s comparable to sonnet, but requires a little bit more instruction and clarity to get things right. For bigger complex changes I still use anthropic’s family, but for very concise and well defined smaller tasks the price of GLM-4.7 is hard to beat.
When left to its own devices, GLM-4.7 frequently tries to build the world. It's also less capable at figuring out stumbling blocks on its own without spiralling.
For small, well-defined tasks, it's broadly comparable to Sonnet.
Given how incredibly cheap it is, it's useful even as a secondary model.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#295Earlier quoted context omitted.
> Going from GLM-4.7 to something comparable to 4.5 or 5.2 would be an absolutely crazy improvement. Before you get too excited, GLM-4.7 outperformed Opus 4.5 on some benchmarks too - https://www.cerebras.ai/blog/glm-4-7 See the LiveCodeBench comparison The benchmarks of the open weights models are always more impressive than the performance. Everyone is competing for attention and market share so the incentives to b…
Sure. My sole point is that calling Opus 4.5 and GPT-5.2 "last generation models" is discounting how good they are. In fact, in my experience, Opus 4.6 isn't much of an improvement over 4.5 for agentic coding. I'm not immediately discounting Z.ai's claims because they showed with GLM-4.7 that they can do quite a lot with very little. And Kimi K2.5 is genuinely a great model, so it's possible for Chinese open-weight m…
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#296The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…
I notice the people who endlessly praise closed-source models never actually USE open weight models, or assume their drop-in prompting methods and workflow will just work for other model families. Especially true for SWEs who used Claude Code first and now think every other model is horrible because they're ONLY used to prompting Claude. It's quite scary to see how people develop this level of worship for a proprietary product that is openly distrusting of users. I am not saying this is true or not of the parent poster, but something I notice in general.
As someone who uses GLM-4.7 a good bit, it's easily at Sonnet 4.5 tier - have not tried GLM-5 but it would be surprising if it wasn't at Opus 4.5 level given the massive parameter increase.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#297Been using GLM-4.7 for a couple weeks now. Anecdotally, it’s comparable to sonnet, but requires a little bit more instruction and clarity to get things right. For bigger complex changes I still use anthropic’s family, but for very concise and well defined smaller tasks the price of GLM-4.7 is hard to beat.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#298Pelican generated via OpenRouter: https://gist.github.com/simonw/cc4ca7815ae82562e89a9fdd99f07... Solid bird, not a great bicycle frame.
How many pelican riding bicycle SVGs were there before this test existed? What if the training data is being polluted with all these wonky results...
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#299Betting on whether they can actually perform their sold behaviors.
Passing around code repositories for years without ever trying to run them, factory sealed.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#300It might be impressive on benchmarks, but there's just no way for them to break through the noise from the frontier models. At these prices they're just hemorrhaging money. I can't see a path forward for the smaller companies in this space.
It's a big deal that open-source capability is less than a year behind frontier models.
And I'm very, very glad it is. A world in which LLM technology is exclusive and proprietary to three companies from the same country is not a good world.