Live data from Hacker News

GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

z.ai

211–220 of 540 posts

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#211

Here is the pricing per M tokens. https://docs.z.ai/guides/overview/pricing Why is GLM 5 more expensive than GLM 4.7 even when using sparse attention? There is also a GLM 5-code model.

It's roughly three times cheaper than GPT-5.2-codex, which in turn reflects the difference in energy cost between US and China.

1. electricity costs are at most 25% of inference costs so even if electricity is 3x cheaper in china that would only be a 16% cost reduction.

2. cost is only a singular input into price determination and we really have absolutely zero idea what the margins on inference even are so assuming the current pricing is actually connected to costs is suspect.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#212
post #149

Earlier quoted context omitted.

> which begs the question: what else is censored or outright changed intentionally? So like every other frontier model that has post training to add safeguards in accordance with local norms. Claude won't help you hotwire a car. Gemini won't write you erotic novels. GPT won't talk about suicide or piracy. etc etc >This is a classic test It's a gotcha question with basic zero real world relevance I'd prefer models to…

The problem with censorship isn't that it degrades performance. The problem is that if the censorship is unilaterally dictated by a government then it becomes a tool for suppression, especially as people use AI more and more for their primary source of information. A company might choose to avoid erotica because it clashes with their brand, or avoid certain topics because they're worried about causing harms. That is…

I'm certainly not in favour of censorship, it just strikes me as silly that it's the first thing people "test" as if it's some cunning insight. Anyone not living under a rock knows tiananmen is censored in anything chinese

>That is very different than centralized

I guess? If the government's modus operandi is the key thing for you when you get access to a new model then yeah maybe it's not for you.

I personally find the western closed model centralised under megacorps model far more alarming, but when a new opus gets released I don't run to tell everyone on hn that I've discovered the new Opus isn't open weight. That would just be silly...

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#213
post #53

Earlier quoted context omitted.

> doesn't make financial sense to self-host I guess that's debatable. I regularly run out of quota on my claude max subscription. When that happens, I can sort of kind of get by with my modest setup (2x RTX3090) and quantized Qwen3. And this does not even account for privacy and availability. I'm in Canada, and as the US is slowly consumed by its spiral of self-destruction, I fully expect at some point a digital iron…

I think AI may be the only place you could get away with calling a 2x350W GPU rig "modest". That's like ten normal computers worth of power for the GPUs alone.

> That's like ten normal computers worth of power for the GPUs alone.

Maybe if your "computer" in question is a smartphone? Remember that the M3 Ultra is a 300w+ chip that won't beat one of those 3090s in compute or raster efficiency.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#216
post #93

Earlier quoted context omitted.

I don't see it as selectable my side either (opencode & max plan)

They updated it now

No luck here. Did you do anything specific to make it show / reauth or something?

ah nvm - found the guidance on how to change it

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#217

The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…

> Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly some benchmaxxing going on. Agreed. I think the problem is that while they can innovate at algorithms and training efficiency, the human part of RLHF just doesn't scale and they can't afford the massive amount of custom data created a…

the new meta is purchasing rl environments where models can be self-corrected (e.g. a compiler will error) after sft + rlhf ran into diminishing returns. although theres still lots of demand for "real world" data for actually economically valuable tasks

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#218

Earlier quoted context omitted.

> doesn't make financial sense to self-host I guess that's debatable. I regularly run out of quota on my claude max subscription. When that happens, I can sort of kind of get by with my modest setup (2x RTX3090) and quantized Qwen3. And this does not even account for privacy and availability. I'm in Canada, and as the US is slowly consumed by its spiral of self-destruction, I fully expect at some point a digital iron…

> I regularly run out of quota on my claude max subscription. When that happens, I can sort of kind of get by with my modest setup (2x RTX3090) and quantized Qwen3. When talking about fallback from Claude plans, The correct financial comparison would be the same model hosted on OpenRouter. You could buy a lot of tokens for the price of a pair of 3090s and a machine to run them.

> You could buy a lot of tokens for the price of a pair of 3090s and a machine to run them.

That's a subjective opinion, to which the answer is "no you can't" for many people.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#219
post #5

It's looking like we'll have Chinese OSS to thank for being able to host our own intelligence, free from the whims of proprietary megacorps. I know it doesn't make financial sense to self-host given how cheap OSS inference APIs are now, but it's comforting not being beholden to anyone or requiring a persistent internet connection for on-premise intelligence. Didn't expect to go back to macOS but they're basically the…

AFAIK they haven't released this one as OSS yet. They might eventually but its pretty obvious to me that at one point all/most those more powerful chinese models probably will stop being OSS.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#220
post #53

Earlier quoted context omitted.

I think AI may be the only place you could get away with calling a 2x350W GPU rig "modest". That's like ten normal computers worth of power for the GPUs alone.

> That's like ten normal computers worth of power for the GPUs alone. Maybe if your "computer" in question is a smartphone? Remember that the M3 Ultra is a 300w+ chip that won't beat one of those 3090s in compute or raster efficiency.

I wouldn't class the M3 Ultra as a "normal" computer either. That's a big-ass workstation. I was thinking along the lines of a typical Macbook or Mac Mini or Windows laptop, which are fine for 99% of anyone who isn't looking to play games or run gigantic AI models locally.
Post reply on HN