Earlier quoted context omitted.
> doesn't make financial sense to self-host I guess that's debatable. I regularly run out of quota on my claude max subscription. When that happens, I can sort of kind of get by with my modest setup (2x RTX3090) and quantized Qwen3. And this does not even account for privacy and availability. I'm in Canada, and as the US is slowly consumed by its spiral of self-destruction, I fully expect at some point a digital iron…
I think AI may be the only place you could get away with calling a 2x350W GPU rig "modest". That's like ten normal computers worth of power for the GPUs alone.
GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
201–210 of 540 posts
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#202Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#203Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#204The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…
Agreed. I think the problem is that while they can innovate at algorithms and training efficiency, the human part of RLHF just doesn't scale and they can't afford the massive amount of custom data created and purchased by the frontier labs.
IIRC it was the application of RLHF which solved a lot of the broken syntax generated by LLMs like unbalanced braces and I still see lots of these little problems in every open source model I try. I don't think I've seen broken syntax from the frontier models in over a year from Codex or Claude.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#205Here is the pricing per M tokens. https://docs.z.ai/guides/overview/pricing Why is GLM 5 more expensive than GLM 4.7 even when using sparse attention? There is also a GLM 5-code model.
It's roughly three times cheaper than GPT-5.2-codex, which in turn reflects the difference in energy cost between US and China.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#206The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…
something that is at parity with Opus 4.5 can ship everything you did in the last 8 weeks, ya know... when 4.5 came out
just remember to put all of this in perspective, most of the engineers and people here haven't even noticed any of this stuff and if they have are too stubborn or policy constrained to use it - and the open source nature of the GLM series helps the policy constrained organizations since they can theoretically run it internally or on prem.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#207It might be impressive on benchmarks, but there's just no way for them to break through the noise from the frontier models. At these prices they're just hemorrhaging money. I can't see a path forward for the smaller companies in this space.
>China’s philosophy is different. They believe model capabilities do not matter as much as application. What matters is how you use AI.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#208Earlier quoted context omitted.
> In 2 years you could make that investment back in pay raise. you can't be a happy uber driver making more money in the next 24 months by having a fancy car fitted with the best FSD in town when all cars in your town have the same FSD.
But they don't have the same human in the loop though.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#209Earlier quoted context omitted.
> doesn't make financial sense to self-host I guess that's debatable. I regularly run out of quota on my claude max subscription. When that happens, I can sort of kind of get by with my modest setup (2x RTX3090) and quantized Qwen3. And this does not even account for privacy and availability. I'm in Canada, and as the US is slowly consumed by its spiral of self-destruction, I fully expect at some point a digital iron…
Your $5,000 PC with 2 GPUs could have bought you 2 years of Claude Max, a model much more powerful and with longer context. In 2 years you could make that investment back in pay raise.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#210The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…
> Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly some benchmaxxing going on. Agreed. I think the problem is that while they can innovate at algorithms and training efficiency, the human part of RLHF just doesn't scale and they can't afford the massive amount of custom data created a…