> 320B total parameters and just 18B active parameters This is pretty hefty for a "flash" model, even a 256 GB setup is insufficient at q4 - and q4 is already the worst-but-still-acceptable quant in my experience. The benchmarks look great, especially since GLM tends to be more honest than the average Chinese lab, but you’ll need to splurge to run it at home. @edit: so many releases that I forgot to math. This fits j…
Speaking as someone who isn't really well versed in this, does 18B active parameters mean that you could potentially hold only the 18B parameters in RAM and stream the rest from a fast NVMe SSD for acceptable performance similar to how Colibri works? https://github.com/JustVugg/colibri
GLM-5.3-Flash
401–410 of 605 posts
Re: GLM-5.3-Flash
#402Earlier quoted context omitted.
I don't understand how people don't consider this. Plus you're spec'd out of near-SOTA level in months. The only reasons to actually do this are a) you have a lot of dispensable income and are a hobbyist/tinkerer, b) you have real, legitimate privacy concerns or, relatedly, c) you're doing something you don't want to get flagged
Not everything is about pure cost. Maybe I don't want to sell my soul supporting the frontier labs because they are straight up pure evil?
Re: GLM-5.3-Flash
#403Re: GLM-5.3-Flash
#404Earlier quoted context omitted.
I've had the exact opposite experience. I've been using 3.8 for my daily driver since last week, and I've gradually been giving it more and more complex tasks as it continues to deliver high quality results. Now I am basically handing off large complex features, and 3.8 is doing the planning, task breakdown, implementation and review with just a few notes from my side. The tradeoff is time (especially on RDMA4 hardwa…
Fascinating that after years of for me this for me that, literally no one writes 2-3 other words like “i do react frontend” or whatever just for us to know why the results are different
Re: GLM-5.3-Flash
#405Earlier quoted context omitted.
What they don’t do. They claim that the GDPR applies if you provide your service in the EU, and that’s a valid claim.
They claim it applies to EU citizens when both they and the service are outside the EU as well.
Re: GLM-5.3-Flash
#406Earlier quoted context omitted.
That's only half the reason it's expensive. The other reason is that it would likely take years to spend $4000 (plus the real cost of electricity) worth of tokens on a 3rd-party provider that's running a similar limited, DS Flash type model. By that time, the hardware will be obsolete, assuming it's still operational.
> it would likely take years to spend $4000 (plus the real cost of electricity) Since that cluster only yields 20-30 tok/s on that size of model, at least a decade before the hardware breaks-even with current token costs, and that's not counting electricity. Assuming continued downward pressure on token prices, and the cost of electricity, it never pays for itself.
Exactly; its a development box for fiddling with GPU hardware with a large amount of video-addressable memory. It's not an inference box, really, though it's neat that I can at all!
Re: GLM-5.3-Flash
#407Earlier quoted context omitted.
I am pretty confident that given a $200 subscription on any of the big labs, you're getting $4000-$8000 per month in subsidized tokens... do what you wan't with your dough... and I too have a spark that I got really early (October 2025), but no, economically it does not compare to what's runnable locally in terms of quality from the frontier models. Economically, it looks like for as long as there are subscriber plan…
This should be obvious but with a model running on local hardware you can do your own RLHF and mod its behavior however you see fit. With cloud hosted models you can't. A few years ago when the models were smaller there were people undoing the guardrails, censorship, and general lobotomization with some form of a RLHF training. You can't do that on larger models unless you have the hardware like this person does. Not…
I do set up my initial runs and likes like quantisation-aware-distillation on my Spark-like to test it out and get it working, so it has value! But its not "worth" it other than its fun hardware to tinker with, IMO.
Re: GLM-5.3-Flash
#408Earlier quoted context omitted.
And don't forget the coolest part, DeepSeek, Qwen, Z.ai and Moonshot have almost caught up while being open about their research and their model weights. We can mostly speculate about OAI and Anthropic models, nothing else, how fun huh?
It would be great if open source US AI companies could get going already.
Re: GLM-5.3-Flash
#409Earlier quoted context omitted.
I have exactly the same opinion Over the last couple years I’ve had to learn sales and understand the thought process behind this better, and I think I’m beginning to understand it The psychology is that most people aren’t really trying to optimize for productivity (even most people who think they are) on an ROI basis, because their compensation is too decoupled from their actual raw output, and more closely coupled…
These are great points. It's a little off topic but what you bring up is why i advise new grads to spend the first couple years of their career in small eat-what-you-kill companies. I think software devs who start out in large companies get this distorted view that their twice a month direct deposit is just magic and comes from the ether no matter what they do. The whole industry would be better off if everyone start…
Re: GLM-5.3-Flash
#410You guys read Z.ai's terms of service, right? Broad and perpetual license over inputs and outputs, and even your name and profile picture. Vague prohibitions on whatever may harm Z.ai’s "interests" or even the "national interests" of any country. Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is. Vague prohibitions on discussing Z.ai, even my posting this comment violates it. Can ban you…
> Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is. I have prompted out a lot of disturbing and inappropriate content with GLM-5.2, that would have left other American models blanched in the face or clutch their pearls. I think this is mostly a reference to Anti-CCP stuff. In fact, I don't think I've ever even had a prompt refused.
I very much have. I've gotten GLM-5.2 refusals for extremely benign security testing on my own infrastructure of the same flavor that people were getting (wrongly) flagged for on Fable during the initial release.