Live data from Hacker News

GLM-5.1: Towards Long-Horizon Tasks

z.ai

191–200 of 285 posts

Re: GLM-5.1: Towards Long-Horizon Tasks

#191
post #118

Every single day, three things are becoming more and more clear: (1) OpenAI & Anthropic are absolutely cooked; it's obvious they have no moat (2) Local/private inference is the future of AI (3) There's *still* no killer product yet (so get to work!)

>(1) OpenAI & Anthropic are absolutely cooked; it's obvious they have no moat

I think big corporations will continue to use them no matter how cheap and good other models are. There's a saying: nobody was fired for buying IBM.

Re: GLM-5.1: Towards Long-Horizon Tasks

#192

Earlier quoted context omitted.

This has got to be bait.. 1) OpenAI and Anthropic are killing it, and continue to do so, their coding tools are unmatched for professionals. 2) Local models don't hold a candle to SOTA models and there's nothing on the horizon that indicates that consumers will be able to run anything close to what you can get in a data center. 3) Coding is a killer product, OpenAI and Anthropic are raking in the cash. The top 3 apps…

The grandparent is definitely wrong on (3). Yes, coding is a killer product, I agree with you. On (2), I agree with you for local models. BUT , there are also the open source Chinese models accessible via open-router. Your argument ("don't hold a candle to SOTA models") does not hold if the comparison is between those. On (1), I agree more with the grandparent than with your assessment. Yes, OpenAI and Anthropic are…

>BUT, there are also the open source Chinese models accessible via open-router.

I thought so myself, but after burning a lot of money on OpenRouter in a few days I just subscribed to Z.ai's Coding Pro plan and using the subscription is much, much friendlier with my wallet.

Re: GLM-5.1: Towards Long-Horizon Tasks

#193
post #177

Earlier quoted context omitted.

The grandparent is definitely wrong on (3). Yes, coding is a killer product, I agree with you. On (2), I agree with you for local models. BUT , there are also the open source Chinese models accessible via open-router. Your argument ("don't hold a candle to SOTA models") does not hold if the comparison is between those. On (1), I agree more with the grandparent than with your assessment. Yes, OpenAI and Anthropic are…

> the open source Chinese models accessible via open-router And? They aren't as good as SOTA models. Even the SOTA model provider's small models aren't worth using for many of my coding tasks.

In my limited experience with it, GLM 5.1 is on par with Opus 4.6.

Re: GLM-5.1: Towards Long-Horizon Tasks

#194
post #156

Earlier quoted context omitted.

> Top-tier models will never run on desktop machines Sorry, but you don't know that

I mean it's not hard to understand that if good model can run on consumer hardware, even better models can run in data centers

Larger, yes, absolutely. Better? Right now it seems that bigger is better, but if we are thinking about long term future, it's not obvious that there isn't a point of diminishing returns with regards to size. I can also imagine a breakthrough, where models become much smaller, with the same or better capabilities as the current, very large ones.

Re: GLM-5.1: Towards Long-Horizon Tasks

#196
Z.ai and their GLM models are pretty low quality.

I've been testing it for awhile now since it seemed to have potential as a local model.

With this new update it still cannot parse simple, test PDFs correctly. It inconsistently tells me that the value in the name field in the document is incorrect, and has the name reversed to put the last name first. Or that a date is wrong as it's in the past/future, when it is not. Tons of fundamental errors like that.

Even when looking at the thinking process there are issues:

I used a test website for it to analyze and it says that the sites copyright year states 2026 which is in the future and to investigate as it could be an attack, but right after prints today's correct date.

I'm in the process of trying to get it uncensored. Hopefully that will create some use out of z.ai

Edit: by the way, which is the best uncensored model at the moment?

Re: GLM-5.1: Towards Long-Horizon Tasks

#197
post #196

Z.ai and their GLM models are pretty low quality. I've been testing it for awhile now since it seemed to have potential as a local model. With this new update it still cannot parse simple, test PDFs correctly. It inconsistently tells me that the value in the name field in the document is incorrect, and has the name reversed to put the last name first. Or that a date is wrong as it's in the past/future, when it is not…

Completely agree with this statement "Z.ai and their GLM models are pretty low quality." I have been trying out and it's kind of useless compare to SOTA models.

Re: GLM-5.1: Towards Long-Horizon Tasks

#198
post #197
post #196

Z.ai and their GLM models are pretty low quality. I've been testing it for awhile now since it seemed to have potential as a local model. With this new update it still cannot parse simple, test PDFs correctly. It inconsistently tells me that the value in the name field in the document is incorrect, and has the name reversed to put the last name first. Or that a date is wrong as it's in the past/future, when it is not…

Completely agree with this statement "Z.ai and their GLM models are pretty low quality." I have been trying out and it's kind of useless compare to SOTA models.

[flagged]

Re: GLM-5.1: Towards Long-Horizon Tasks

#200
post #20

To be honest I am a bit sad as, glm5.1 is producing mich better typescript than opus or codex imo, but no matter what it does sometimes go into shizo mode at some point over longer contexts. Not always tho I have had multiple session go over 200k and be fine.

When it works and its not slow it can impress. Like yesterday it solved something that kimi k2.5 could not. and kimi was best open source model for me. But it still slow sometimes. I have z.ai and kimi subscription when i run out of tokens for claude (max) and codex(plus). i have a feeling its nearing opus 4.5 level if they could fix it getting crazy after like 100k tokens.

Why don't you start a new session or use the /compact command when context gets to 100k tokens?

From my testing it was ok until 145k tokens, the largest context I had before switching to a new session. I think Z.ai officially said it should be good until 200k tokens.

Using it in Open Code is compacting the context automatically when it gets too large.

Post reply on HN