Live data from Hacker News

GLM-5.1: Towards Long-Horizon Tasks

z.ai

241–250 of 285 posts

Re: GLM-5.1: Towards Long-Horizon Tasks

#241
post #220

Earlier quoted context omitted.

> And to be clear, OpenAI/Anthropic most definitely know this: that's why they've been aquihiring like crazy, trying to find that one team that will make the thing. Anthropic is up to $30B annual recurring revenue. I wish I had failing business models like that. > Token prices are significantly subsidized and anyone that does any serious work with AI can tell you this. Go use an almost-SOTA model (a big Deepseek or Q…

> Anthropic is up to $30B annual recurring revenue. I wish I had failing business models like that. And profit? A company can have $300B annual revenue, and still be a failing business if it's making a loss. Somewhere along the line we seem to have forgotten this basic fact. Eventually there will be no more rounds of funding to feed the fire.

Costs can always be optimized, revenue is much harder to optimize.

Re: GLM-5.1: Towards Long-Horizon Tasks

#242

Earlier quoted context omitted.

I hit this as well. It just seems to hang and process for ages.

Try lowering thinking level with GLM-5.1, to me that seems to have an impact on mitigating the blocking behaviour.

Hmm I'll try that, but OpenCode shows me the thinking and it's not even doing that. I'm just getting no tokens from it at all.

Re: GLM-5.1: Towards Long-Horizon Tasks

#243
post #202
post #196

Z.ai and their GLM models are pretty low quality. I've been testing it for awhile now since it seemed to have potential as a local model. With this new update it still cannot parse simple, test PDFs correctly. It inconsistently tells me that the value in the name field in the document is incorrect, and has the name reversed to put the last name first. Or that a date is wrong as it's in the past/future, when it is not…

I'e been using their models pretty much daily for the past 2 months to work on the codebase of a very complex B2B2C platform written in an unusual functional language (F#) with an angular frontend. I also use Claude premium daily for another client, and i use Codex. and i can tell you that GLM5 is at this point much more capable than Claude and Codex for complex backend end work, complex feature planning, and long ho…

“GLM5…better than Opus, Codex, Gemini…”

What wild claim to make. Unsupported by benchmarks, unsupported by the consensus of the community, no evidence provided.

Sounds like in another comment here even the GLM5 team concedes they are behind the frontier wrt tool calling, do you know something they don’t?

Re: GLM-5.1: Towards Long-Horizon Tasks

#245
post #224
post #196

Z.ai and their GLM models are pretty low quality. I've been testing it for awhile now since it seemed to have potential as a local model. With this new update it still cannot parse simple, test PDFs correctly. It inconsistently tells me that the value in the name field in the document is incorrect, and has the name reversed to put the last name first. Or that a date is wrong as it's in the past/future, when it is not…

I still use GLM 4.7 for well defined coding tasks. I never got 5.0 to work satisfactorily, it felt like a hosting problem (z.ai) where it would work for a while then, for whatever reason, it couldn't respond to the context any more - but that's just a hunch. I had no such trouble with 4.7 and find it fast and productive. Haven't tried 5.1; am using openAI models for coding most of the time.

Same here.

Z.ai seem to promote 4.7 for smaller tasks, 5.1 for larger tasks (similar to Anthropic's recommendation for usage of Haiku and Sonnet/Opus models).

5.1 works for me already in the most economical basic paid tier ("lite coding plan"), unlike first release of v5 (5.0 ?)

Re: GLM-5.1: Towards Long-Horizon Tasks

#246

Earlier quoted context omitted.

For conversational purposes that may be too slow, but as a coding assistant this should work, especially if many tasks are batched, so that they may progress simultaneously through a single pass over the SSD data.

Three hour coffee break while the LLM prepares scaffolding for the project.

Rather, Imagine you have 2-3 of these working 24/7 on top of what you're doing today. What does your backlog look like a month from now?

Re: GLM-5.1: Towards Long-Horizon Tasks

#247
post #202

Earlier quoted context omitted.

I'e been using their models pretty much daily for the past 2 months to work on the codebase of a very complex B2B2C platform written in an unusual functional language (F#) with an angular frontend. I also use Claude premium daily for another client, and i use Codex. and i can tell you that GLM5 is at this point much more capable than Claude and Codex for complex backend end work, complex feature planning, and long ho…

“GLM5…better than Opus, Codex, Gemini…” What wild claim to make. Unsupported by benchmarks, unsupported by the consensus of the community, no evidence provided. Sounds like in another comment here even the GLM5 team concedes they are behind the frontier wrt tool calling, do you know something they don’t?

I know my use case and my personal experience :) i am not trying to pretend that it is the best in benchmarks, just sharing my experience so people know that some folks are having a very good experience with GLM models, compared to the competition.

My only goal is to encourage people to try it out so they can see if it moves the needle for them, because there are fair chances that it will. I am not trying to start a flamewar or something.

Re: GLM-5.1: Towards Long-Horizon Tasks

#248
post #118

Every single day, three things are becoming more and more clear: (1) OpenAI & Anthropic are absolutely cooked; it's obvious they have no moat (2) Local/private inference is the future of AI (3) There's *still* no killer product yet (so get to work!)

[deleted]

Re: GLM-5.1: Towards Long-Horizon Tasks

#249
post #182

Earlier quoted context omitted.

I don't want to respond to 100 comments about the same thing, and this one happens to be on top, so, in my humble opinion: (1): You don't have to be an Ed Zitron disciple to infer that OpenAI and Anthropic are likely overvalued and that Nvidia is selling everyone shovels in a gold rush. AI is a game-changing technology, but a shitty chat interface does not a company make. OpenAI and Anthropic need to recoup astronomi…

> Go use an almost-SOTA model (a big Deepseek or Qwen model) offered by many bare-metal providers and you'll see what "true" token prices should look like. Qwen3.5-122B-A10B is $0.26 input, $2.08 output. Where's the subsidy? It's ten times cheaper than Opus. Or did you mean that we're subsidizing their training? But then "OpenClaw token-vomit on top of Claude is fiscally untenable" makes no sense. Yeah, I don't know…

[deleted]

Re: GLM-5.1: Towards Long-Horizon Tasks

#250
post #182

Earlier quoted context omitted.

I don't want to respond to 100 comments about the same thing, and this one happens to be on top, so, in my humble opinion: (1): You don't have to be an Ed Zitron disciple to infer that OpenAI and Anthropic are likely overvalued and that Nvidia is selling everyone shovels in a gold rush. AI is a game-changing technology, but a shitty chat interface does not a company make. OpenAI and Anthropic need to recoup astronomi…

> Go use an almost-SOTA model (a big Deepseek or Qwen model) offered by many bare-metal providers and you'll see what "true" token prices should look like. Qwen3.5-122B-A10B is $0.26 input, $2.08 output. Where's the subsidy? It's ten times cheaper than Opus. Or did you mean that we're subsidizing their training? But then "OpenClaw token-vomit on top of Claude is fiscally untenable" makes no sense. Yeah, I don't know…

Maybe he's comparing the renting price of a bare metal server on its own, and doesn't realise how drastically cheaper they are to batch together for an API provider.
Post reply on HN