Live data from Hacker News

GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

z.ai

141–150 of 540 posts

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#141
post #132

Earlier quoted context omitted.

Apple devices have high memory bandwidth necessary to run LLMs at reasonable rates. It’s possible to build a Linux box that does the same but you’ll be spending a lot more to get there. With Apple, a $500 Mac Mini has memory bandwidth that you just can’t get anywhere else for the price.

> a $500 Mac Mini has memory bandwidth that you just can’t get anywhere else for the price. The cheapest new mac mini is $600 on Apple's US store. And it has a 128-bit memory interface using LPDDR5X/7500, nothing exotic. The laptop I bought last year for <$500 has roughly the same memory speed and new machines are even faster.

> The cheapest new mac mini is $600 on Apple's US store.

And you're only getting 16GB at that base spec. It's $1000 for 32GB, or $2000 for 64GB plus the requisite SOC upgrade.

> And it has a 128-bit memory interface using LPDDR5X/7500, nothing exotic.

Yeah, 128-bit is table stakes and AMD is making 256-bit SOCs as well now. Apple's higher end Max/Ultra chips are the ones which stand out with their 512 and 1024-bit interfaces. Those have no direct competition.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#142

[flagged]

As I promised earlier: https://news.ycombinator.com/item?id=46781777 "I will save this for the future, when people complain about Chinese open models and tell me: But this Chinese LLM doesn't respond to question about Tianmen square." Please stop using Tianmen question as an example to evaluate the company or their models: https://news.ycombinator.com/item?id=46779809

Neither should be censoring objective reality.

Why defend it on either side?

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#143
post #79

Do we know if it as vision? That is lacking from 4.7, you need to use an mcp for it.

It does not have vision. On the Z.ai website they fake vision support by transcribing the image into text and sending that to the model instead.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#144
post #58

Earlier quoted context omitted.

> doesn't make financial sense to self-host I guess that's debatable. I regularly run out of quota on my claude max subscription. When that happens, I can sort of kind of get by with my modest setup (2x RTX3090) and quantized Qwen3. And this does not even account for privacy and availability. I'm in Canada, and as the US is slowly consumed by its spiral of self-destruction, I fully expect at some point a digital iron…

Did the napkin math on M3 Ultra ROI when DeepSeek V3 launched: at $0.70/2M tokens and 30 tps, a $10K M3 Ultra would take ~30 years of non-stop inference to break even - without even factoring in electricity. Clearly people aren't self-hosting to save money. I've got a lite GLM sub $72/yr which would require 138 years to burn through the $10K M3 Ultra sticker price. Even GLM's highest cost Max tier (20x lite) at $720/…

I don't think an Apple PC can run full Deepseek or GLM models.

Even if you quantize the hell out of the models to fit in the memory, they will be very slow.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#145

Soft launch? I can't find a blog post on their website.

There was a one-line X post about something new being available at their chat endpoint, but that's about it at the time of this writing. Nothing at GitHub or HuggingFace, no tech report or anything.

What's funny is it's available on /v1/models, but if you call it you get an error saying it's not accessible yet. No word on pricing, probably the same as 4.7 if I had to guess (0.6/2.2)

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#146
post #3

Wut? Was glm 4.7 not just a few weeks ago? I wonder if I will be able to use it with my coding plan. Paid just 9 usd for 3 month.

What's the use case for Zai/GLM? I'm currently on Claude Pro, and the Zai looks about 50% more expensive after the first 3 months and according to their chart GLM 4.7 is not quite as capable as Opus 4.5? I'm looking to save on costs because I use it so infrequently, but PAYG seems like it'd cost me more in a single session per month than the monthly cost plan.

> What's the use case for Zai/GLM?

It's cheap :) It seems they stopped it now, but for the last 2 month you could buy the lite plan for a whole year for under 30 USD, while claude is ~19 USD per month. I bought 3 month for ~9 USD.

I use it for hobby projects. Casual coding with Open Code.

If price is not important Opus / Codex are just plain better.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#147
post #5

It's looking like we'll have Chinese OSS to thank for being able to host our own intelligence, free from the whims of proprietary megacorps. I know it doesn't make financial sense to self-host given how cheap OSS inference APIs are now, but it's comforting not being beholden to anyone or requiring a persistent internet connection for on-premise intelligence. Didn't expect to go back to macOS but they're basically the…

Not going to call $30/mo for a github copilot subscription "cheap". More like "extortionary".

Yeah it's funny how the needle has moved on this kind of thing.

Two years ago people scoffed at buying a personal license for e.g. JetBrains IDEs which netted out to $120 USD or something a year; VS Code etc took off because they were "free"

But now they're dumping monthly subs to OpenAI and Anthropic that work out to the same as their car insurance payments.

It's not sustainable.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#148
post #5

It's looking like we'll have Chinese OSS to thank for being able to host our own intelligence, free from the whims of proprietary megacorps. I know it doesn't make financial sense to self-host given how cheap OSS inference APIs are now, but it's comforting not being beholden to anyone or requiring a persistent internet connection for on-premise intelligence. Didn't expect to go back to macOS but they're basically the…

I'm not sure being beholden to the whims of the Chinese Communist Party is an iota better than the whims of proprietary megacorps, especially given this probably will become part of a megacorp anyway.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#149
post #136
post #89

Earlier quoted context omitted.

You're surprised that chinese model makers try to follow chinese law?

This is a classic test to see if the model is censored, as censorship is rarely limited to just one event, which begs the question: what else is censored or outright changed intentionally?

> which begs the question: what else is censored or outright changed intentionally?

So like every other frontier model that has post training to add safeguards in accordance with local norms.

Claude won't help you hotwire a car. Gemini won't write you erotic novels. GPT won't talk about suicide or piracy. etc etc

>This is a classic test

It's a gotcha question with basic zero real world relevance

I'd prefer models to be uncensored too because it does harm overall performance but this is such a non-issue in practice

Post reply on HN