Live data from Hacker News

GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

z.ai

181–190 of 540 posts

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#181
The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now.

Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly some benchmaxxing going on.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#186
post #179

why don't they publish at ARC-AGI ? too expensive?

Arc agi was never a good benchmark that tested spatial understanding more than reasoning. I'm glad it's no longer popular

What do you mean? It definitely tests reasoning as well, and if anything, I expect spatial and embodied reasoning to become more important in the coming years, as AI agents will be expected to take on more real world tasks.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#187
It's live on openrouter now.

In my personal benchmark it's bad. So far the benchmark has been a really good indicator of instruction following and agentic behaviour in general.

To those who are curious, the benchmark is just the ability of model to follow a custom tool calling format. I ask it to using coding tasks using chat.md [1] + mcps. And so far it's just not able to follow it at all.

[1] https://github.com/rusiaaman/chat.md

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#189
post #169
post #91

Earlier quoted context omitted.

Your $5,000 PC with 2 GPUs could have bought you 2 years of Claude Max, a model much more powerful and with longer context. In 2 years you could make that investment back in pay raise.

> In 2 years you could make that investment back in pay raise. you can't be a happy uber driver making more money in the next 24 months by having a fancy car fitted with the best FSD in town when all cars in your town have the same FSD.

But they don't have the same human in the loop though.
Post reply on HN