Live data from Hacker News

GLM-5.3 is now open-weight

huggingface.co

221–230 of 296 posts

Re: GLM-5.3 is now open-weight

#221
post #173

Earlier quoted context omitted.

Too hot and expensive to run right now but a great hedge for peace of mind against $200 subscriptions shooting up to the $4000* they should cost. *$1000? $14,000? Who knows but everything in the middle there has been claimed.

If they "should" cost 4k in the sense of marginal cost, then you will be spending more running the same at home, because your home hardware will always be less efficient.

There is a big difference in the cost of a 5-nines up time system in a heavily space constrained environment compared to a home hobby white box used for some coding. The GPUs alone cost 10x for the data center versions compared to the gaming versions even with similar specs.

The cost of online services is also largely a result of the cost of training (though hard to say exactly what that number is). Assuming you are using open weight models at home, you aren't paying for the training - someone else is.

Re: GLM-5.3 is now open-weight

#222
post #29

Earlier quoted context omitted.

on my TrustedRouter: z-ai/glm-5.3: also Z.ai, Novita, Atlas Cloud, IO.NET

How am I supposed to navigate around there? For example the pricing page is empty or is that how it was supposed to look? On models and providers pages there are lists but no way to filter or get any kind of meaningful info. Or is this a WIP/POC?

updated the pricing page link to the models page, and added search

we are doing billions of tokens a day and thousands of users

Re: GLM-5.3 is now open-weight

#224

GLM 5.3 is probably the sweet spot open weights model if you want to go beyond deepseek flash or the new glm flash. I used it with pi and had a fairly good time, especially since it’s less touchy about cyber and whatnot than the US guys. It’s slightly behind Kimi in ability but it’s a lot easier to run it, I’d expect prices (and speed!) from third parties to be noticeably better. Assuming you’re willing to drop a fat…

When we consider: * LLM usage is new for the world * Models are evolving quickly with high worldwide competition * Hardware is evolving despite RAM shortages Is investing a huge sum of money in equipment for local inference a wise use of money? Or are M5 Ultra and equivalently priced local inference hardware future-proof enough to be worth it relative to how the market is evolving? Maybe it’s all a question of what y…

Only reason to spend a bunch of money on hardware to run LLMs locally is if it's a hobby to you to an extent that even renting the GPUs temporarily won't satisfy you.

Re: GLM-5.3 is now open-weight

#227

Earlier quoted context omitted.

What if the model is hopelessly obsolete, and thus no demand, but I want that specific model? Owning the weights and hardware is not just solving for one problem. It eliminates all the classes of problems that occur outside of your building, if you have a solar and battery setup. Also, on a more practical basis, what if the way it's served is bad. Maybe I want my specific KV setup, or ultra low quant for entertaining…

> What if the model is hopelessly obsolete, and thus no demand, but I want that specific model? You can still find a lot of old and completely outdated models on OpenRouter. The providers can scale serving of models up and down as demand arrives, so models don't generally disappear. They're just kept in the mix and the clouds will allocate hardware to it if someone is willing to pay. In the odd case that it disappear…

> There is a certain personality type that cannot tolerate any uncertainty and must lock everything in right now against all future possibilities.

I thought HN banned personal attacks. I'm in this sentence and I don't like it. /s

I just buy the good apple hardware because it's good, and it also happens to run local models. It's not as good for the dollar, don't get me wrong, but I'm not going to develop iOS without a mac, that's even more questionable than buying a strix or whatever.

Re: GLM-5.3 is now open-weight

#228
post #221
post #173

Earlier quoted context omitted.

If they "should" cost 4k in the sense of marginal cost, then you will be spending more running the same at home, because your home hardware will always be less efficient.

There is a big difference in the cost of a 5-nines up time system in a heavily space constrained environment compared to a home hobby white box used for some coding. The GPUs alone cost 10x for the data center versions compared to the gaming versions even with similar specs. The cost of online services is also largely a result of the cost of training (though hard to say exactly what that number is). Assuming you are…

> The cost of online services is also largely a result of the cost of training

OpenRouter prices are somewhat simmilar to Antrhopic/OpenAI API prices. So I conclude that the hardware plus operating margin alone can genuinely produce prices way above what you'd pay if you had a subscription. Of course the primary unkown factor is average token use per subscription. Without that it's all wild speculation.

Re: GLM-5.3 is now open-weight

#229

Earlier quoted context omitted.

I think the biggest reason is to own the stack so your model can't be changed out from under you, but maybe I care about that too much.

> I think the biggest reason is to own the stack so your model can't be changed out from under you, The concern would be future regulations that prohibit you from buying a hosted version of the model. Even that could be bypassed with a VPN to another country but it's more work to go through the payments. As long as there is demand for a model, it will be hosted by multiple providers.

The only reason I'm considering picking one up is I think we're not that far away from compute limitations in consumer hardware.

Re: GLM-5.3 is now open-weight

#230

Earlier quoted context omitted.

When we consider: * LLM usage is new for the world * Models are evolving quickly with high worldwide competition * Hardware is evolving despite RAM shortages Is investing a huge sum of money in equipment for local inference a wise use of money? Or are M5 Ultra and equivalently priced local inference hardware future-proof enough to be worth it relative to how the market is evolving? Maybe it’s all a question of what y…

I have a Strix Halo and dual 32GB GPUs in my desktop, that sit idle right now, because the electricity to run them and to cool them in 110F weather Texas is currently experiencing pretty much nulls any savings I might see over getting better models from cloud providers. While I mostly use Claude or Codex with subscriptions for agentic work, for API use DeepSeek has usually been my go to, but now I guess it's GLM 5.3…

I wish Texas would write up a regulation allowing 'balcony solar' as I could easily generate 1000-2000w of solar in my small back yard to take a bite out the sizeable cooling bill I have.
Post reply on HN