Live data from Hacker News

GLM-5.3 is now open-weight

huggingface.co

231–240 of 296 posts

Re: GLM-5.3 is now open-weight

#231

Earlier quoted context omitted.

I think the biggest reason is to own the stack so your model can't be changed out from under you, but maybe I care about that too much.

> I think the biggest reason is to own the stack so your model can't be changed out from under you, The concern would be future regulations that prohibit you from buying a hosted version of the model. Even that could be bypassed with a VPN to another country but it's more work to go through the payments. As long as there is demand for a model, it will be hosted by multiple providers.

I live in a place where using VPN is illegal and akin to "terrorism" because why would you want to hide what you are doing. Only bad guys hide. So if you use VPN, you are a bad guy.

https://srinagar.nic.in/notice/immediate-suspension-of-virtu...

Phones are randomly searched on the streets and if VPN is found, arrested

https://www.medianama.com/2026/01/223-jammu-kashmir-vpn-ban-...

https://timesofindia.indiatimes.com/india/after-vpn-ban-in-k...

“Out of the 15 individuals identified, five were minors who were counselled and advised in the presence of their guardians, with emphasis on awareness, lawful digital conduct, and the consequences of violating lawful orders,” he added.

Re: GLM-5.3 is now open-weight

#232

GLM 5.3 is probably the sweet spot open weights model if you want to go beyond deepseek flash or the new glm flash. I used it with pi and had a fairly good time, especially since it’s less touchy about cyber and whatnot than the US guys. It’s slightly behind Kimi in ability but it’s a lot easier to run it, I’d expect prices (and speed!) from third parties to be noticeably better. Assuming you’re willing to drop a fat…

In terms of pure tokens per dollar, absolutely not worth it. That said, when I bought my pair of Sparks, the best model I could run on it was GPT OSS 120B. That has an AA score of 24. Today, the best model I can run on them is GLM 5.3 Flash at Q4, AA score 57. Just still out on GLM 5.3 mixed quant. So from that perspective, they are many times better value than when I bought them, and will likely continue to increase…

> GLM 5.3 Flash at Q4, AA score 57

That AA score is for the original model only

Re: GLM-5.3 is now open-weight

#233

Earlier quoted context omitted.

Tools vs services in my mind. There is no guarantee any provider will continue to do what they are doing for you at the price they are doing it. The object permanence of not having to reinvent the world every time a model gets sunsetted has value.

> Tools vs services in my mind. There is no guarantee any provider will continue to do what they are doing for you at the price they are doing it. with open models, there is ecosystem/market of providers, where you can easily switch to provider you like

Until there's an executive order that blocks one model from being served.

Re: GLM-5.3 is now open-weight

#234
post #43

Earlier quoted context omitted.

I have just built an Epyc with 512gb DDR4 3200 RAM for a "reasonable" price and I'm hoping to have a setup with GLM as the architect and Qwen 27b/Next Flash as the implementer. This is 1/5 of the price of the Mac, but also probably 1/5 of the speed lol.

I’ll be very curious what you get with DDR4. I also almost went that way. I have an Epyc DDR 5 rig and the best I see is 10 tok/s. Caveat being that’s at Q8 and a 4090 doing pre fill so it could be pushed up. The surprising thing for me is how much work you will need to cool the banks if you’re near your memory ceiling. My memory starts soft throttling at about 74C (dies may be hotter, that’s the bank temp) and will…

Typically computers with these larger memory amounts have fans that scream like a banshee trying to move impossible amounts of air over the memory and CPU. Getting something both cool and quite can be a bit difficult.

Re: GLM-5.3 is now open-weight

#235
post #230

Earlier quoted context omitted.

I have a Strix Halo and dual 32GB GPUs in my desktop, that sit idle right now, because the electricity to run them and to cool them in 110F weather Texas is currently experiencing pretty much nulls any savings I might see over getting better models from cloud providers. While I mostly use Claude or Codex with subscriptions for agentic work, for API use DeepSeek has usually been my go to, but now I guess it's GLM 5.3…

I wish Texas would write up a regulation allowing 'balcony solar' as I could easily generate 1000-2000w of solar in my small back yard to take a bite out the sizeable cooling bill I have.

Seems like it's easier to ask forgiveness than permission. And, I wouldn't bet on this legislature ever doing anything that would disempower fossil energy or reduce their profits, even a little bit.

Re: GLM-5.3 is now open-weight

#236
Model is great. Have to try with oh my py --prewalk with either Kimi as planner and GLM 5.3 as implementer or GLM 5.3 as planner and Flash as implementer.

They have a new shiny data center with Chinese GPUs, so hopefully they will be able to handle the demand. Their subscriptions are meh, in particular the Flash model has not that much usage. The other thing I wanted to try is Groq Qwen 3.8-27B at 450 tps to see if it's able to get work done faster.

I feel it's very important to play outside of the walled gardens, I feel they will pull the rug very soon, and I need to get job done while waiting hardware prices to go down...

Re: GLM-5.3 is now open-weight

#237
post #75

Earlier quoted context omitted.

When we consider: * LLM usage is new for the world * Models are evolving quickly with high worldwide competition * Hardware is evolving despite RAM shortages Is investing a huge sum of money in equipment for local inference a wise use of money? Or are M5 Ultra and equivalently priced local inference hardware future-proof enough to be worth it relative to how the market is evolving? Maybe it’s all a question of what y…

It is absolutely not worth buying hardware to run models for purely (long term) cost reasons. For open weights models the economies of scale means the cloud beats local significantly and your payback time is like 10 years. However there are other reasons (e.g. privacy) that might make it worth running locally for some people.

I'm actively uninspired to write high quality code when using Anthropic/OpenAI models given the high chance I'm a customer as well as used as dataset generation tool for them.

But currently cloud does beat costs of hardware ownership, particularly with ridiculously high RAM/GPU/SSD costs....again due to these same companies.

Re: GLM-5.3 is now open-weight

#238

Earlier quoted context omitted.

With competition we kind of have guarantee up to what providers can do, they don't have that much control, the most radical thing they can do is to go bankrupt.

Have you already forgotten the Fable drama that happened just two months ago?

Yes, I remember recurring extensions on plan inclusion then becoming permanent - that's my point.

Re: GLM-5.3 is now open-weight

#239
post #224

Earlier quoted context omitted.

When we consider: * LLM usage is new for the world * Models are evolving quickly with high worldwide competition * Hardware is evolving despite RAM shortages Is investing a huge sum of money in equipment for local inference a wise use of money? Or are M5 Ultra and equivalently priced local inference hardware future-proof enough to be worth it relative to how the market is evolving? Maybe it’s all a question of what y…

Only reason to spend a bunch of money on hardware to run LLMs locally is if it's a hobby to you to an extent that even renting the GPUs temporarily won't satisfy you.

Or if you need stuff that APIs don't / can't provide. Or for future proofing your workflows. Running things locally gets you "the same thing" in perpetuity, while APIs might change, models can be deprecated and features can be removed.

Cybersec is also hit and miss, depending on what provider you choose, verification systems and all that jazz. Also, running locally allows you 100% data privacy, in any situation and for whatever usecase you might have. ~100k for hardware for a small team of devs to code locally is not that expensive in the grand scheme of things.

Lastly, local models allow for training / finetuning on your own data and processes. $/tok is not everything for everyone. Sometimes you can take a hit on value / speed if you get something else that matters for you.

Re: GLM-5.3 is now open-weight

#240

Earlier quoted context omitted.

In terms of pure tokens per dollar, absolutely not worth it. That said, when I bought my pair of Sparks, the best model I could run on it was GPT OSS 120B. That has an AA score of 24. Today, the best model I can run on them is GLM 5.3 Flash at Q4, AA score 57. Just still out on GLM 5.3 mixed quant. So from that perspective, they are many times better value than when I bought them, and will likely continue to increase…

> GLM 5.3 Flash at Q4, AA score 57 That AA score is for the original model only

Then take DeepSeek V4 flash with AA score 52. Runs unquantized on 2x DGX spark with 1M context.
Post reply on HN