Live data from Hacker News

GLM-5.3 is now open-weight

huggingface.co

151–160 of 296 posts

Re: GLM-5.3 is now open-weight

#151

Earlier quoted context omitted.

So far I don’t regret buying an M1 Max device with 32Gb of RAM. The models available for it keep getting better (running just about okay for interactive use) and 400 GB/s of bandwidth is still considered a lot. The models are currently improving much faster than the hardware and this doesn’t seem to have plateaued yet.

Cool! I'm thinking about a local set up. What's your usual tokens/second rate?

Not OP, but I’m running local models on a M1 Max as well with 64GB RAM.

It varies by model, but I’m getting 50-60 t/s with Qwen 3.6 35B and Qwen 3 coder 30B.

I’ve also used Qwen 3.8 27B but I get 10t/s on it.

It’s useable in some use cases, but I rely mostly on my $20 Claude subscription.

Re: GLM-5.3 is now open-weight

#152
post #91

Does this mean it'll be on Bedrock soon? I hear great things about this model but I want AWS data handling practices...

I doubt it - AWS hasn't added any non-western models since GLM 5 and MiniMax M2.5 in February, afaik. Might be a deal with OpenAI (GPT 5.4 was the first to be available via Bedrock, in April) or might just be that there isn't a lot of demand due to corporate skittishness around models trained in China.

many of them (kimi k3, glm-5.3) have license requirements to sell them with model-as-a-service.

Re: GLM-5.3 is now open-weight

#153

Earlier quoted context omitted.

Tools vs services in my mind. There is no guarantee any provider will continue to do what they are doing for you at the price they are doing it. The object permanence of not having to reinvent the world every time a model gets sunsetted has value.

With competition we kind of have guarantee up to what providers can do, they don't have that much control, the most radical thing they can do is to go bankrupt.

Have you already forgotten the Fable drama that happened just two months ago?

Re: GLM-5.3 is now open-weight

#154
post #47

Earlier quoted context omitted.

I think the region is just the HQ of the provider. So z.ai's region is Singapore but it's quite likely that their servers are actually in China

I don't think that's right, or if it is, OpenRouter has incorrect data. Several Chinese companies (headquartered in China) have Singapore listed as their region on OR. And some companies, like Alibaba Cloud, have multiple regions listed. I'm happy to be proven wrong, but this makes me think that the region is where the servers are, not where the HQ is.

I couldn't find any article that states z.ai has a data center in Singapore. There are stories of their new 1 GW data center in China though. Also, openrouter lists HQs on their providers page which matches the regions. https://openrouter.ai/providers

Re: GLM-5.3 is now open-weight

#155

I previously posted that DS4Flash was _good_ but not _great_ on two DGX Sparks, but I have to say that GLM-5.3 is pretty amazing. It's been able to tackle all the random hard problems I've thrown at it and it has the intuition that DS4Flash seems to lack. We're nowhere near a Fable-class model IMO, but things are going to get interesting in this next year.

It's early days, but GLM 5.3 Flash is the first local model that feels good enough to me to be a "main" model without debating whether each problem needs to be sent to a stronger model. DS4 Flash is good enough at implementing given a plan, but I wasn't always a fan of what it came up with when asked to plan something.

The good news is it can only get better from here.

Re: GLM-5.3 is now open-weight

#156
post #140

Earlier quoted context omitted.

When we consider: * LLM usage is new for the world * Models are evolving quickly with high worldwide competition * Hardware is evolving despite RAM shortages Is investing a huge sum of money in equipment for local inference a wise use of money? Or are M5 Ultra and equivalently priced local inference hardware future-proof enough to be worth it relative to how the market is evolving? Maybe it’s all a question of what y…

Jalapeno is matching or very near Vera Rubin at 1/4 the power. I would not buy hardware now.

OpenAI have only just announced it and have every reason to hype it up.

Could be a long time till gets released

Re: GLM-5.3 is now open-weight

#157
post #43

Earlier quoted context omitted.

I have just built an Epyc with 512gb DDR4 3200 RAM for a "reasonable" price and I'm hoping to have a setup with GLM as the architect and Qwen 27b/Next Flash as the implementer. This is 1/5 of the price of the Mac, but also probably 1/5 of the speed lol.

I’ll be very curious what you get with DDR4. I also almost went that way. I have an Epyc DDR 5 rig and the best I see is 10 tok/s. Caveat being that’s at Q8 and a 4090 doing pre fill so it could be pushed up. The surprising thing for me is how much work you will need to cool the banks if you’re near your memory ceiling. My memory starts soft throttling at about 74C (dies may be hotter, that’s the bank temp) and will…

I am getting 10t/s on unsloth's Q3kxl with 2x3090s@250w. It's enough for me for now. I will probably upgrade the GPUs down the line. DDR5 would have made the price of the machine double and I just wasn't prepared to pay that much.

Temp wise, no throttling, surprisingly cool.

Re: GLM-5.3 is now open-weight

#158
post #43

Earlier quoted context omitted.

I have just built an Epyc with 512gb DDR4 3200 RAM for a "reasonable" price and I'm hoping to have a setup with GLM as the architect and Qwen 27b/Next Flash as the implementer. This is 1/5 of the price of the Mac, but also probably 1/5 of the speed lol.

Curious about that price, if you don't mind sharing a ballpark

About 5k with RAM and GPUs bought used. Eastern Europe.

Re: GLM-5.3 is now open-weight

#159
post #43

GLM 5.3 is probably the sweet spot open weights model if you want to go beyond deepseek flash or the new glm flash. I used it with pi and had a fairly good time, especially since it’s less touchy about cyber and whatnot than the US guys. It’s slightly behind Kimi in ability but it’s a lot easier to run it, I’d expect prices (and speed!) from third parties to be noticeably better. Assuming you’re willing to drop a fat…

I have just built an Epyc with 512gb DDR4 3200 RAM for a "reasonable" price and I'm hoping to have a setup with GLM as the architect and Qwen 27b/Next Flash as the implementer. This is 1/5 of the price of the Mac, but also probably 1/5 of the speed lol.

It’s not unified ram? I.e VRAM so it will struggle

Re: GLM-5.3 is now open-weight

#160

GLM 5.3 is probably the sweet spot open weights model if you want to go beyond deepseek flash or the new glm flash. I used it with pi and had a fairly good time, especially since it’s less touchy about cyber and whatnot than the US guys. It’s slightly behind Kimi in ability but it’s a lot easier to run it, I’d expect prices (and speed!) from third parties to be noticeably better. Assuming you’re willing to drop a fat…

When we consider: * LLM usage is new for the world * Models are evolving quickly with high worldwide competition * Hardware is evolving despite RAM shortages Is investing a huge sum of money in equipment for local inference a wise use of money? Or are M5 Ultra and equivalently priced local inference hardware future-proof enough to be worth it relative to how the market is evolving? Maybe it’s all a question of what y…

I have a Strix Halo and dual 32GB GPUs in my desktop, that sit idle right now, because the electricity to run them and to cool them in 110F weather Texas is currently experiencing pretty much nulls any savings I might see over getting better models from cloud providers. While I mostly use Claude or Codex with subscriptions for agentic work, for API use DeepSeek has usually been my go to, but now I guess it's GLM 5.3 or the Flash version. And, for security work that Anthropic or OpenAI models are likely to refuse, I've been using Kimi K3 (also via subscription, though their subscription is extremely stingy), but I guess GLM is now the one for that, too.

Anyway, yeah, even at the prices I spent on my local AI stuff (I bought before RAMpocalypse really kicked into gear, so I bought old server GPUs for about $350 each and the Strix Halo for a little over $2k) it was never going to pay for itself; I just like to tinker. But, I can't imagine spending today's prices for hardware for local AI.

When the memory shortage ends, I'll be down to the Apple Store (or, more likely, clicking refresh on the Apple outlet every few days). But, until then, there continues to be a glut of cheap and free models in the cloud that are better than anything I can run locally and they're faster, too.

Post reply on HN