Live data from Hacker News

GLM-5.1: Towards Long-Horizon Tasks

z.ai

231–240 of 285 posts

Re: GLM-5.1: Towards Long-Horizon Tasks

#232
post #220
post #182

Earlier quoted context omitted.

I don't want to respond to 100 comments about the same thing, and this one happens to be on top, so, in my humble opinion: (1): You don't have to be an Ed Zitron disciple to infer that OpenAI and Anthropic are likely overvalued and that Nvidia is selling everyone shovels in a gold rush. AI is a game-changing technology, but a shitty chat interface does not a company make. OpenAI and Anthropic need to recoup astronomi…

> And to be clear, OpenAI/Anthropic most definitely know this: that's why they've been aquihiring like crazy, trying to find that one team that will make the thing. Anthropic is up to $30B annual recurring revenue. I wish I had failing business models like that. > Token prices are significantly subsidized and anyone that does any serious work with AI can tell you this. Go use an almost-SOTA model (a big Deepseek or Q…

> Anthropic is up to $30B annual recurring revenue. I wish I had failing business models like that.

And profit? A company can have $300B annual revenue, and still be a failing business if it's making a loss.

Somewhere along the line we seem to have forgotten this basic fact. Eventually there will be no more rounds of funding to feed the fire.

Re: GLM-5.1: Towards Long-Horizon Tasks

#233

Earlier quoted context omitted.

No killer product? Coding assistants and LLM's in general are the single most awe-inspiring achievement of humanity in my lifetime, technological or otherwise. They've already massively improved my and others' lives and they're only going to get better. If pre and post industrial revolution used to be the major binary delineation of our history, I'm fairly confident it will soon be seen as pre and post AI instead.

> Coding assistants and LLM's in general are the single most awe-inspiring achievement of humanity in my lifetime Landing a man on the moon is way more impressive. Finding several vaccines for a once in a century pandemic within a year of its outbreak is and achievement that in its impact and importance dwarfs what the entire LLM industry put together has achieved. The near-complete eradication of polio, once again,…

Those are all good things, but with the current AI boom we've invented something with the potential to invent those kinds of things on its own, if not now then in the near future. It's far more important and impactful to invent a digital mind that can invent an arbitrary number of vaccines than to just invent one vaccine, no matter how hard it was to invent the vaccine by hand.

Re: GLM-5.1: Towards Long-Horizon Tasks

#235
post #221

Earlier quoted context omitted.

Is there really no rule that discourages 99% of your interactions with HN from being peddling some useless slop benchmark?

If it's relevant to the discussion, I hope not. I've spent probably over100 hours working on this benchmarking/site platform, and all tests are manually written. For me (and many others that reached out to me) are not useless either. I use this myself regularly when choosing and comparing new models. I honestly beleive it is providing value to the conversation. Let me know if you know of a better platform you can use…

It's a great benchmark. Don't listen to the haters. This one is especially interesting.

https://aibenchy.com/compare/anthropic-claude-sonnet-4-6-med...

Re: GLM-5.1: Towards Long-Horizon Tasks

#237
post #220
post #182

Earlier quoted context omitted.

I don't want to respond to 100 comments about the same thing, and this one happens to be on top, so, in my humble opinion: (1): You don't have to be an Ed Zitron disciple to infer that OpenAI and Anthropic are likely overvalued and that Nvidia is selling everyone shovels in a gold rush. AI is a game-changing technology, but a shitty chat interface does not a company make. OpenAI and Anthropic need to recoup astronomi…

> And to be clear, OpenAI/Anthropic most definitely know this: that's why they've been aquihiring like crazy, trying to find that one team that will make the thing. Anthropic is up to $30B annual recurring revenue. I wish I had failing business models like that. > Token prices are significantly subsidized and anyone that does any serious work with AI can tell you this. Go use an almost-SOTA model (a big Deepseek or Q…

It is easy to get 30B when you resell something you buy for 50B

Re: GLM-5.1: Towards Long-Horizon Tasks

#238
post #196

Z.ai and their GLM models are pretty low quality. I've been testing it for awhile now since it seemed to have potential as a local model. With this new update it still cannot parse simple, test PDFs correctly. It inconsistently tells me that the value in the name field in the document is incorrect, and has the name reversed to put the last name first. Or that a date is wrong as it's in the past/future, when it is not…

From what I gather qwen is currently the undisputed local LLM king.

Re: GLM-5.1: Towards Long-Horizon Tasks

#239
post #182

Earlier quoted context omitted.

This has got to be bait.. 1) OpenAI and Anthropic are killing it, and continue to do so, their coding tools are unmatched for professionals. 2) Local models don't hold a candle to SOTA models and there's nothing on the horizon that indicates that consumers will be able to run anything close to what you can get in a data center. 3) Coding is a killer product, OpenAI and Anthropic are raking in the cash. The top 3 apps…

I don't want to respond to 100 comments about the same thing, and this one happens to be on top, so, in my humble opinion: (1): You don't have to be an Ed Zitron disciple to infer that OpenAI and Anthropic are likely overvalued and that Nvidia is selling everyone shovels in a gold rush. AI is a game-changing technology, but a shitty chat interface does not a company make. OpenAI and Anthropic need to recoup astronomi…

> Go use an almost-SOTA model (a big Deepseek or Qwen model) offered by many bare-metal providers and you'll see what "true" token prices should look like.

Qwen3.5-122B-A10B is $0.26 input, $2.08 output. Where's the subsidy? It's ten times cheaper than Opus. Or did you mean that we're subsidizing their training? But then "OpenClaw token-vomit on top of Claude is fiscally untenable" makes no sense.

Yeah, I don't know where you got your costs from. Bare metal providers are significantly cheaper than Anthropic.

Re: GLM-5.1: Towards Long-Horizon Tasks

#240
post #224

Earlier quoted context omitted.

I still use GLM 4.7 for well defined coding tasks. I never got 5.0 to work satisfactorily, it felt like a hosting problem (z.ai) where it would work for a while then, for whatever reason, it couldn't respond to the context any more - but that's just a hunch. I had no such trouble with 4.7 and find it fast and productive. Haven't tried 5.1; am using openAI models for coding most of the time.

I hit this as well. It just seems to hang and process for ages.

Try lowering thinking level with GLM-5.1, to me that seems to have an impact on mitigating the blocking behaviour.
Post reply on HN