Live data from Hacker News

GLM-5.3 is now open-weight

huggingface.co

51–60 of 296 posts

Re: GLM-5.3 is now open-weight

#51
post #43

GLM 5.3 is probably the sweet spot open weights model if you want to go beyond deepseek flash or the new glm flash. I used it with pi and had a fairly good time, especially since it’s less touchy about cyber and whatnot than the US guys. It’s slightly behind Kimi in ability but it’s a lot easier to run it, I’d expect prices (and speed!) from third parties to be noticeably better. Assuming you’re willing to drop a fat…

I have just built an Epyc with 512gb DDR4 3200 RAM for a "reasonable" price and I'm hoping to have a setup with GLM as the architect and Qwen 27b/Next Flash as the implementer. This is 1/5 of the price of the Mac, but also probably 1/5 of the speed lol.

Honestly I suspect neither of them will be performing terribly well but with DDR4 3200 RAM I wonder if you'll be counting tokens per second or seconds per token. I mean, you do at least get a lot of memory channels at least, compared to consumer PCs. I am curious to hear what performance you get, I feel there is not enough information out there on what different setups manage to eek out.

Re: GLM-5.3 is now open-weight

#52
post #2

I've been using it more and more. Feels like Opus 4.8, in the best possible way.

I really like how it doesn't have that Claude talk. It just does the thing without Claude's "load-bearing honesty." It's probably my favorite model to interact with, even if it isn't the best or most reliable.

My second favorite model by now is GLM 5.3 flash which is very capable of day to day task. I use it as the main model and GLM 5.3 for task that is more complex

Re: GLM-5.3 is now open-weight

#53
post #51
post #43

Earlier quoted context omitted.

I have just built an Epyc with 512gb DDR4 3200 RAM for a "reasonable" price and I'm hoping to have a setup with GLM as the architect and Qwen 27b/Next Flash as the implementer. This is 1/5 of the price of the Mac, but also probably 1/5 of the speed lol.

Honestly I suspect neither of them will be performing terribly well but with DDR4 3200 RAM I wonder if you'll be counting tokens per second or seconds per token. I mean, you do at least get a lot of memory channels at least, compared to consumer PCs. I am curious to hear what performance you get, I feel there is not enough information out there on what different setups manage to eek out.

The fastest I was able to get my Threadripper 3960X + 2x 3090s + 256GB DDR4-3200 to run a 2-bit quant of GLM-5.2 was 8 TPS. I would expect to be in seconds-per-token territory for a pure-CPU 4-bit quant.

Re: GLM-5.3 is now open-weight

#54

How much usage do you find you get on these kinda models (I know the pricing changes a bit) compared to a $20 sub say for Google AI Pro in anti gravity? I hate how difficult it is to compare prices when looking at subscriptions. Would $20 in open router, using models like GLM get me more or less?

I think it'd get you less than a $20 sub to any of the big three. I've used it on OpenRouter and found it kind of expensive for the results, but that might change now that it's open weight and other providers can host it/compete with Z.ai. For the work I did with it, I would've rather used DeepSeek V4 Flash just because it's more economical and still gives good results IMO.

Z.ai does have their own subscription, but I haven't used it because their privacy policy was pretty buns last time I checked.

Re: GLM-5.3 is now open-weight

#55
post #39
post #13

Earlier quoted context omitted.

It’s cheaper sure, but it’s very slow. It’s not a drop in replacement

I think we don't have a good draft model for better speculative decoding yet (e.g. DFlash 2). Once we do, it will be faster.

It very well could be faster, but right now it isn’t.

Re: GLM-5.3 is now open-weight

#56
post #36
post #25

I'd like to ask Sam Altman if he still thinks that it's too dangerous to publish GPT-3. I mean, no one would use it, but what is his reasoning for not publishing it now, in 2026?

They already publish gpt-oss which is several generations better than gpt-3

GPT-3 is a different model than gpt-oss and is therefore not an answer to the question.

I cannot stand using gpt-oss, but I miss some of the creative spark of GPT-3 davinci dearly.

Re: GLM-5.3 is now open-weight

#57
post #36

Earlier quoted context omitted.

They already publish gpt-oss which is several generations better than gpt-3

At release, GPT-OSS was arguably a few generations behind the open frontier.

No? They were the frontier, or near it, at the time of release: https://artificialanalysis.ai/models/releases/gpt-oss-120b

Re: GLM-5.3 is now open-weight

#58

Earlier quoted context omitted.

One could also run it locally on a used dual xeon (or amd-equivalent) server with 512GB RAM, albeit slower, if you have a useful workflow for it that's like "take this day's efforts and run it through various analysis agents", combined with giving it one-shot tasks/modules to build overnight. You would want a place like a garage or basement to put the server because it'll be loud.

You'd also likely spend far more in electricity than the API cost of processing the prompt(s)

yes, though for some uses, not sending data anywhere to third parties has its own value which is harder to measure.

Re: GLM-5.3 is now open-weight

#59
post #8

Earlier quoted context omitted.

It's actually slightly more expensive ($0.50 vs $0.48), but there's a temporary 50% discount. I've seen dozens of conversations about it in last 24 hours, and every major inference provided added in first 24 hours. I think it's gaining plenty of traction.

It's interesting that OpenCode Go is treating it as 2x more expensive than DeepSeek Flash, even factoring in the 50% discount

OpenCode Go is becoming less of a good deal by the month. I pretty much only use it for mimo 2.5 pro now, and everything else is either ollama or openrouter.

Re: GLM-5.3 is now open-weight

#60
post #43

GLM 5.3 is probably the sweet spot open weights model if you want to go beyond deepseek flash or the new glm flash. I used it with pi and had a fairly good time, especially since it’s less touchy about cyber and whatnot than the US guys. It’s slightly behind Kimi in ability but it’s a lot easier to run it, I’d expect prices (and speed!) from third parties to be noticeably better. Assuming you’re willing to drop a fat…

I have just built an Epyc with 512gb DDR4 3200 RAM for a "reasonable" price and I'm hoping to have a setup with GLM as the architect and Qwen 27b/Next Flash as the implementer. This is 1/5 of the price of the Mac, but also probably 1/5 of the speed lol.

Depending on which Epyc you got it might be slower than 1/5 of the speed.
Post reply on HN