Live data from Hacker News

GLM-5.3 is now open-weight

huggingface.co

161–170 of 296 posts

Re: GLM-5.3 is now open-weight

#162
post #51
post #43

Earlier quoted context omitted.

I have just built an Epyc with 512gb DDR4 3200 RAM for a "reasonable" price and I'm hoping to have a setup with GLM as the architect and Qwen 27b/Next Flash as the implementer. This is 1/5 of the price of the Mac, but also probably 1/5 of the speed lol.

Honestly I suspect neither of them will be performing terribly well but with DDR4 3200 RAM I wonder if you'll be counting tokens per second or seconds per token. I mean, you do at least get a lot of memory channels at least, compared to consumer PCs. I am curious to hear what performance you get, I feel there is not enough information out there on what different setups manage to eek out.

What model are you interested in? DS Flash 0731@Q4KXL I'm about 25-30tps. Same as the new Qwen3.8 Flash Next. The new GLM 5.3Q3KXL at 10tps. I've got 2x3090s which I didn't mention in the original message.

Re: GLM-5.3 is now open-weight

#163
post #159
post #43

Earlier quoted context omitted.

I have just built an Epyc with 512gb DDR4 3200 RAM for a "reasonable" price and I'm hoping to have a setup with GLM as the architect and Qwen 27b/Next Flash as the implementer. This is 1/5 of the price of the Mac, but also probably 1/5 of the speed lol.

It’s not unified ram? I.e VRAM so it will struggle

I'm getting about 10tps @Q3kxl with 2x3090s.

Re: GLM-5.3 is now open-weight

#164
post #60
post #43

Earlier quoted context omitted.

I have just built an Epyc with 512gb DDR4 3200 RAM for a "reasonable" price and I'm hoping to have a setup with GLM as the architect and Qwen 27b/Next Flash as the implementer. This is 1/5 of the price of the Mac, but also probably 1/5 of the speed lol.

Depending on which Epyc you got it might be slower than 1/5 of the speed.

48c 7643. I'm getting about 10tps @Q3kxl with 2x3090s.

Re: GLM-5.3 is now open-weight

#165
What's very promising here is the number of tokens-vs-accuracy ratio. I am assuming their "output tokens" means tokens generated as part of thinking and any tool calls (what are referred to as "input tokens" from billing PoV by service providers). The Chinese models like Qwen3.8 and GLM 5.2 are insanely overthinking in our workloads (which are highly complex data analysis tasks). It's a factor of 3-4x over Opus and GPT models. Even with cheaper prices per 1M tokens, the cost ends up being higher, in some cases 2x. So this is very promising from GLM 5.3. Looking forward to trying it.

Re: GLM-5.3 is now open-weight

#168

Is it possible to fine tune this model and unlock / extend its cybersecurity capabilities? I'm scared that maybe we are not ready for an open-weight model with high cybersecurity skills.

I sure as hell hope so. I want to point these models at my own computers and harden everything I own.

The insufferable gatekeeping of the US companies is actively contributing to computer insecurity at this point.

Re: GLM-5.3 is now open-weight

#169
post #91

Does this mean it'll be on Bedrock soon? I hear great things about this model but I want AWS data handling practices...

What data handling practices exactly? If it's about privacy there is TensorX.ai, which claim to host in Europe and be GDPR compliant

Re: GLM-5.3 is now open-weight

#170

Earlier quoted context omitted.

I think it'd get you less than a $20 sub to any of the big three. I've used it on OpenRouter and found it kind of expensive for the results, but that might change now that it's open weight and other providers can host it/compete with Z.ai. For the work I did with it, I would've rather used DeepSeek V4 Flash just because it's more economical and still gives good results IMO. Z.ai does have their own subscription, but…

> but I haven't used it because their privacy policy was pretty buns last time I checked. What did you find objectionable? I looked at it when I subscribed almost a year ago and I was fine with it (e.g. they don't train on your API inputs).

Z.ai gives itself a perpetual license to all of your inputs and outputs.
Post reply on HN