Live data from Hacker News

GLM-5.1: Towards Long-Horizon Tasks

z.ai

251–260 of 285 posts

Re: GLM-5.1: Towards Long-Horizon Tasks

#251
post #220

Earlier quoted context omitted.

> And to be clear, OpenAI/Anthropic most definitely know this: that's why they've been aquihiring like crazy, trying to find that one team that will make the thing. Anthropic is up to $30B annual recurring revenue. I wish I had failing business models like that. > Token prices are significantly subsidized and anyone that does any serious work with AI can tell you this. Go use an almost-SOTA model (a big Deepseek or Q…

It is easy to get 30B when you resell something you buy for 50B

The proverbial "50B" is investment in next year's model. The current model cost under "30B", and therefore "is profitable". It is a bet on scaling, yes, but that's been common throughout the industry (see, eg, Amazon not being profitable for many years but building infrastructure)

Re: GLM-5.1: Towards Long-Horizon Tasks

#252
post #251

Earlier quoted context omitted.

It is easy to get 30B when you resell something you buy for 50B

The proverbial "50B" is investment in next year's model. The current model cost under "30B", and therefore "is profitable". It is a bet on scaling, yes, but that's been common throughout the industry (see, eg, Amazon not being profitable for many years but building infrastructure)

Except the rumors are they subsidize even the inference, not that they have capex in training.

Re: GLM-5.1: Towards Long-Horizon Tasks

#253
post #220

Earlier quoted context omitted.

> And to be clear, OpenAI/Anthropic most definitely know this: that's why they've been aquihiring like crazy, trying to find that one team that will make the thing. Anthropic is up to $30B annual recurring revenue. I wish I had failing business models like that. > Token prices are significantly subsidized and anyone that does any serious work with AI can tell you this. Go use an almost-SOTA model (a big Deepseek or Q…

> Anthropic is up to $30B annual recurring revenue. I wish I had failing business models like that. And profit? A company can have $300B annual revenue, and still be a failing business if it's making a loss. Somewhere along the line we seem to have forgotten this basic fact. Eventually there will be no more rounds of funding to feed the fire.

Anthropic has raised $64B in total since they were founded.

Even if you say we are going to measure profit in the very special hacker news way of looking at money taken in from customer revenue against money invested and we say they can't do things like counting building data centers or buying GPUs as capital expenses and instead have to count them against profit then in 2 years time they will have made more money than they have taken in investment.

That is extraordinary.

Re: GLM-5.1: Towards Long-Horizon Tasks

#254
post #251

Earlier quoted context omitted.

The proverbial "50B" is investment in next year's model. The current model cost under "30B", and therefore "is profitable". It is a bet on scaling, yes, but that's been common throughout the industry (see, eg, Amazon not being profitable for many years but building infrastructure)

Except the rumors are they subsidize even the inference, not that they have capex in training.

The maths shows inference is very profitable. Look at how Google/AWS/Azure change the same rates as Anthropic does for running Claude models.

Re: GLM-5.1: Towards Long-Horizon Tasks

#255
post #177

Earlier quoted context omitted.

> the open source Chinese models accessible via open-router And? They aren't as good as SOTA models. Even the SOTA model provider's small models aren't worth using for many of my coding tasks.

In my limited experience with it, GLM 5.1 is on par with Opus 4.6.

I used GLM5 quite a bit, and I'd say it was maybe on par with Sonnet for most simple to medium tasks. Definitely not Opus though. Didn't test super long context tasks, and that's where I would expect it to break down. A recent study on software maintainability still showed Sonnet and Opus were peerless on that metric, although GLM series of models has been making impressive gains.

Re: GLM-5.1: Towards Long-Horizon Tasks

#256

Earlier quoted context omitted.

This has got to be bait.. 1) OpenAI and Anthropic are killing it, and continue to do so, their coding tools are unmatched for professionals. 2) Local models don't hold a candle to SOTA models and there's nothing on the horizon that indicates that consumers will be able to run anything close to what you can get in a data center. 3) Coding is a killer product, OpenAI and Anthropic are raking in the cash. The top 3 apps…

The grandparent is definitely wrong on (3). Yes, coding is a killer product, I agree with you. On (2), I agree with you for local models. BUT , there are also the open source Chinese models accessible via open-router. Your argument ("don't hold a candle to SOTA models") does not hold if the comparison is between those. On (1), I agree more with the grandparent than with your assessment. Yes, OpenAI and Anthropic are…

Open models are good but if you need a $10k GPU to run them then 99% of people are better of subscribing to OAI or CC.

Nowadays I also feel model performance matters less than the design of the tool harness, inference speed, and the other systems that surround a typical coding model.

Re: GLM-5.1: Towards Long-Horizon Tasks

#257
post #196

Z.ai and their GLM models are pretty low quality. I've been testing it for awhile now since it seemed to have potential as a local model. With this new update it still cannot parse simple, test PDFs correctly. It inconsistently tells me that the value in the name field in the document is incorrect, and has the name reversed to put the last name first. Or that a date is wrong as it's in the past/future, when it is not…

>by the way, which is the best uncensored model at the moment? There are no such models, depending on your definition of censorship. If you're referring to abliteration and similar automated techniques, they're snake oil.

That is absolutely not the case. Try HauHauCS's Qwen 3.5 models. They don't refuse anything, and they don't lose a noticeable amount of capability.

Re: GLM-5.1: Towards Long-Horizon Tasks

#259
post #105

Not only did this one draw me an excellent pelican... it also animated it! https://simonwillison.net/2026/Apr/7/glm-51/

Surely at this point it’s part of the training set and the benchmark has lost its value?

these comments are as useless as simon posting his pelicans
Post reply on HN