Live data from Hacker News

GLM-5.1: Towards Long-Horizon Tasks

z.ai

201–210 of 285 posts

Re: GLM-5.1: Towards Long-Horizon Tasks

#201

We're still adding samples, but some early takeaways from benchmarking on https://gertlabs.com : Contrary to the model card, its one-shot performance is more impressive than its agentic abilities. On both metrics, GLM 5.1 is competitive with frontier models. But keeping in mind this is an open source model operating near the frontier, it's nothing short of incredible. I suspect 2 issues with the model are keeping it…

It would be nice if you can test the model with different harnesses, Z.ai's own Z Code, Claude Code, Open Code, Pi, Cursor etc. My impression is that the choice of harness matters a lot.

Interesting idea. The metric I'd intuitively want to see is low variance between harnesses for a smarter model. But if a large sample of models statistically outperformed with a certain harness, that's indeed a valuable signal for a developer.

Re: GLM-5.1: Towards Long-Horizon Tasks

#202
post #196

Z.ai and their GLM models are pretty low quality. I've been testing it for awhile now since it seemed to have potential as a local model. With this new update it still cannot parse simple, test PDFs correctly. It inconsistently tells me that the value in the name field in the document is incorrect, and has the name reversed to put the last name first. Or that a date is wrong as it's in the past/future, when it is not…

I'e been using their models pretty much daily for the past 2 months to work on the codebase of a very complex B2B2C platform written in an unusual functional language (F#) with an angular frontend.

I also use Claude premium daily for another client, and i use Codex. and i can tell you that GLM5 is at this point much more capable than Claude and Codex for complex backend end work, complex feature planning, and long horizon tasks. One thing i've noticed is that it is particularly good at following instructions and guidelines, even deep into the execution of a plan.

To me the only problem is that z.ai have had trouble with inference : the performance of their API has been pretty poor at times. It looks like this is an hardware issue related to the Huawei chips they use rather than an issue with the model itself. The situation has been substantially improving over the past few weeks.

GLM5.1, GLM5-Turbo and GLM5v are at this point better than Opus, Codex, Gemini and other claude source models. We have reached a major turning point. To me, the only closed source model still in the game is codex as it is much faster at executing simple tasks and implementing already created plans.

Try GLM5v for your PDF work, it's their last generation vision model that has been released a couple of days ago.

Re: GLM-5.1: Towards Long-Horizon Tasks

#203
post #202
post #196

Z.ai and their GLM models are pretty low quality. I've been testing it for awhile now since it seemed to have potential as a local model. With this new update it still cannot parse simple, test PDFs correctly. It inconsistently tells me that the value in the name field in the document is incorrect, and has the name reversed to put the last name first. Or that a date is wrong as it's in the past/future, when it is not…

I'e been using their models pretty much daily for the past 2 months to work on the codebase of a very complex B2B2C platform written in an unusual functional language (F#) with an angular frontend. I also use Claude premium daily for another client, and i use Codex. and i can tell you that GLM5 is at this point much more capable than Claude and Codex for complex backend end work, complex feature planning, and long ho…

[flagged]

Re: GLM-5.1: Towards Long-Horizon Tasks

#204
post #203
post #202

Earlier quoted context omitted.

I'e been using their models pretty much daily for the past 2 months to work on the codebase of a very complex B2B2C platform written in an unusual functional language (F#) with an angular frontend. I also use Claude premium daily for another client, and i use Codex. and i can tell you that GLM5 is at this point much more capable than Claude and Codex for complex backend end work, complex feature planning, and long ho…

[flagged]

I appreciate that it's not working for your use case but it's unfortunate that you dismiss the experience of others. And i am not chinese, I am European. Thanks for your feedback anyway.

Re: GLM-5.1: Towards Long-Horizon Tasks

#205
post #118

Every single day, three things are becoming more and more clear: (1) OpenAI & Anthropic are absolutely cooked; it's obvious they have no moat (2) Local/private inference is the future of AI (3) There's *still* no killer product yet (so get to work!)

I was trying to use Claude.ai today to learn how to do hexagonal geometry. Every time I asked a question it generated an interactive geometry graph on the fly in Javascript. Sometimes it spent minutes compiling and testing code on the server so it could make sure it was correct. I was really impressed. Anyway I couldn't really learn anything since when the code didn't work I wasn't sure if I had ported it wrong or th…

This is also my exact experience

Re: GLM-5.1: Towards Long-Horizon Tasks

#206
post #203
post #202

Earlier quoted context omitted.

I'e been using their models pretty much daily for the past 2 months to work on the codebase of a very complex B2B2C platform written in an unusual functional language (F#) with an angular frontend. I also use Claude premium daily for another client, and i use Codex. and i can tell you that GLM5 is at this point much more capable than Claude and Codex for complex backend end work, complex feature planning, and long ho…

[flagged]

Sounds like you two are taking pass each other. PDF work is a specific niche that according to you it fails, the other person say it's good at coding.

Re: GLM-5.1: Towards Long-Horizon Tasks

#207
post #206
post #203

Earlier quoted context omitted.

[flagged]

Sounds like you two are taking pass each other. PDF work is a specific niche that according to you it fails, the other person say it's good at coding.

Scroll down to my other comment, I've used it specifically for coding as well.

"It couldn't even debug some moderately complicated python scripts reliably."

Re: GLM-5.1: Towards Long-Horizon Tasks

#209
post #203
post #202

Earlier quoted context omitted.

I'e been using their models pretty much daily for the past 2 months to work on the codebase of a very complex B2B2C platform written in an unusual functional language (F#) with an angular frontend. I also use Claude premium daily for another client, and i use Codex. and i can tell you that GLM5 is at this point much more capable than Claude and Codex for complex backend end work, complex feature planning, and long ho…

[flagged]

I tried Gemini 3.1 pro once to implement a previously designed 7-phase plan. it only implemented a quarter of the plan before stopping, the code didnt even compile because half of the scaffolding was missing. it then confidently said everything was done.

Codex and GLM didnt have any issue following the exact same plan and getting a working app. So I would argue Gemini is the failure here.

Re: GLM-5.1: Towards Long-Horizon Tasks

#210
post #118

Every single day, three things are becoming more and more clear: (1) OpenAI & Anthropic are absolutely cooked; it's obvious they have no moat (2) Local/private inference is the future of AI (3) There's *still* no killer product yet (so get to work!)

No killer product? Coding assistants and LLM's in general are the single most awe-inspiring achievement of humanity in my lifetime, technological or otherwise. They've already massively improved my and others' lives and they're only going to get better. If pre and post industrial revolution used to be the major binary delineation of our history, I'm fairly confident it will soon be seen as pre and post AI instead.

> Coding assistants and LLM's in general are the single most awe-inspiring achievement of humanity in my lifetime

Landing a man on the moon is way more impressive. Finding several vaccines for a once in a century pandemic within a year of its outbreak is and achievement that in its impact and importance dwarfs what the entire LLM industry put together has achieved. The near-complete eradication of polio, once again, way more important and impactful.

Post reply on HN