Live data from Hacker News

GLM-5.1: Towards Long-Horizon Tasks

z.ai

111–120 of 285 posts

Re: GLM-5.1: Towards Long-Horizon Tasks

#112
A bit off-topic, but for some reason, even though I don't use LLMs for my job or for my hobbies, or in daily life frequently (and when I do, it's mostly some kind of "rubber duck brainstorm"), when I see open-weight releases like this one or the recent Gemma 4 (which is very good for local models); the first time was with DeepSeek-R1 (this one, despite being blamed for "censorship", was heavily censored only via DeepSeek API, the local model - full-weight 685B, not the distilled ones - was pretty much unhinged regarding censorship on any topic)... there's always one song coming to mind and I simply can't get rid of it no matter how hard I try.

"I am the storm that is approaching, provoking..." : )

Re: GLM-5.1: Towards Long-Horizon Tasks

#113
post #81

One of the bench maxed models . Every time I tried it , it’s not on par even with other open source models .

Feeling very much the same. Attempting to use it through Claude Code as a model it just completely lost all context on what it was doing after a few months and kept short circuiting even with the most helpful prompts I could give, outside of just writing out the answer myself. I really do not get the praise for this model.

Being "better than Opus 4.6" is not really something a benchmark will tell you. It's much more a consensus of users liking the flavor of an answer, rather than fueling x% correct on a benchmark.

Re: GLM-5.1: Towards Long-Horizon Tasks

#115
post #114
post #105

Not only did this one draw me an excellent pelican... it also animated it! https://simonwillison.net/2026/Apr/7/glm-51/

Simon, you need to come up with improved benchmarks soon.

Agree. But you can keep the pelican theme in whatever new benchmark you choose to come up with. Iconic at this point.

Re: GLM-5.1: Towards Long-Horizon Tasks

#116

Earlier quoted context omitted.

Three hour coffee break while the LLM prepares scaffolding for the project.

Like computing used to be. When I first compiled a Linux kernel it ran overnight on a Pentium-S. I had little idea what I was doing, probably compiled all the modules by mistake.

At least the compiler was free

Re: GLM-5.1: Towards Long-Horizon Tasks

#117

I am on their "Coding Lite" plan, which I got a lot of use out of for a few months, but it has been seriously gimped now. Obvious quantization issues, going in circles, flipping from X to !X, injecting chinese characters. It is useless now for any serious coding work.

Is there any advantage to their fixed payment plans at all vs just using this model via 3rd party providers via openrouter, given how relatively cheap they tend to be on a per-token basis? Providers like DeepInfra are already giving access to 5.1 https://deepinfra.com/zai-org/GLM-5.1 $1.40 in $4.40 out $0.26 cached / 1M tokens That's more expensive than other models, but not terrible, and will go down over time, and…

I use GLM 5 Turbo sporadically for a client, and my Openrouter expense might climb over a dollar per day if I insist. At about 20 work days per month it's an easy choice.

Re: GLM-5.1: Towards Long-Horizon Tasks

#118
Every single day, three things are becoming more and more clear:

    (1) OpenAI & Anthropic are absolutely cooked; it's obvious they have no moat
    (2) Local/private inference is the future of AI
    (3) There's *still* no killer product yet (so get to work!)

Re: GLM-5.1: Towards Long-Horizon Tasks

#119
post #118

Every single day, three things are becoming more and more clear: (1) OpenAI & Anthropic are absolutely cooked; it's obvious they have no moat (2) Local/private inference is the future of AI (3) There's *still* no killer product yet (so get to work!)

What benefit is there to dropping $50k on GPUs to run this personally besides being a cool enthusiast project?

Re: GLM-5.1: Towards Long-Horizon Tasks

#120
post #118

Every single day, three things are becoming more and more clear: (1) OpenAI & Anthropic are absolutely cooked; it's obvious they have no moat (2) Local/private inference is the future of AI (3) There's *still* no killer product yet (so get to work!)

No killer product? Coding assistants and LLM's in general are the single most awe-inspiring achievement of humanity in my lifetime, technological or otherwise. They've already massively improved my and others' lives and they're only going to get better. If pre and post industrial revolution used to be the major binary delineation of our history, I'm fairly confident it will soon be seen as pre and post AI instead.
Post reply on HN