Live data from Hacker News

GLM-5.2 – How to Run Locally

unsloth.ai

101–110 of 328 posts

Re: GLM-5.2 – How to Run Locally

#101
post #12
post #6

wonder if AMD's new ai chip can run this with ease? I'm seriously consider buying it. GLM 5.2 is just shy of GPT 5.4 so I would welcome offloading any grunt work locally I am very excited for local LLMs I think we may have GPT 5.5-xhigh level of performance for under 2000 EUR This should put more pressure on the frontier models to avoid sitting on any fancy stuff and lower token prices as a whole. Nothing beats a loc…

The AMD 395 supports up to 128GB unified RAM. So still not enough even at 1-bit quant unfortunately.

96gb vram is the max it supports.

Re: GLM-5.2 – How to Run Locally

#102
post #81
post #68

Earlier quoted context omitted.

Crossing my fingers that this boom jumpstarts 90's like improvements in computing hardware. I feel like part of the reason for the relative stagnation in hardware over the last twenty years was simply the lack of use cases to justify hardware refreshes by businesses. Most of the money and energy went to mobile for the last fifteen years. Affordable local inference might be the gravy train the server, desktop, and lap…

>I feel like part of the reason for the relative stagnation in hardware over the last twenty years was simply the lack of use cases to justify hardware refreshes by businesses. No, we're running into limits of moore's law, and it's showing in prices for new nodes, where they're getting denser but not cheaper.

It's true we hit limits, but I feel like a lot of it was "limits" in the sense that the tradeoff stopped being worth the cost, so we optimized in other areas.

So we hit limits on clock speed in the early 2000s (ex - the 4ghz wall) but it also turned out that mobile as the driver for sales meant no one really cared much about clock speed compared to performance/watt.

Clock speed mattered, but only relative to how many watts it took to get it (and above 4ghz... too many watts).

But we've seen a 15x improvement over the last 20 years. Performance/Watt is WAY up.

My guess is that LLMs are going to drive another "improvement cycle" in areas that we didn't care much about before.

I've built about 10 personal desktop machines (1 every ~4 years) and I can honestly say that I didn't care much about memory bandwidth prior to 2021.

In the same way that I didn't care much about how many watts my pentium 4 was using in 2005.

But now... now I care a lot about memory bandwidth. I care about memory speeds and total system ram in a manner I really, really didn't before.

So I think we're going to see a big shift to machines built on unified ram with a crazy focus on squeezing memory bandwidth and total ram capacity as far as we can.

My bet is that we'll get a similar 10-15x improvement by 2040 in unified system ram designs.

I fully expect to see 2tb unified ram desktops and 200gb unified ram phones be relatively common on a 20 year timeline, assuming we see similar levels of geopolitical stability (ex - world war 3 throws a wrench into things).

Re: GLM-5.2 – How to Run Locally

#103
post #32

Earlier quoted context omitted.

The ram/gpu shortage won't last forever though. Moreover we can be pretty confident that long-term the prices will obey wrights law and come down in cost significantly (from the pre-shortage prices) as we learn to produce them more efficiently. LLM companies are valued as if they're going to have some enduring monopoly that they can extract money from... GLM-5.2 and similar models make that valuation very very questi…

> The ram/gpu shortage won't last forever though. No disagreement there, but it could easily last another 3 to 5 years which is a long time in tech terms.

Long enough for them to IPO and all the execs to retire. I doubt they care beyond the IPO.

Re: GLM-5.2 – How to Run Locally

#104

Earlier quoted context omitted.

I suspect the time horizon is shorter because of software advances. We are getting more capability out of smaller models Alibaba released Qwen 3.6 "tiny" models not that long ago, they punch way above their weight(s)

> Alibaba released Qwen 3.6 "tiny" models not that long ago, they punch way above their weight(s) True, Qwen3.6-27B is amazing for it's size. However, it seems likely that we're not going to see anymore of these smaller models from Alibaba/Qwen since several key players exited that organization a few months back.

Good to know, I think the trend is clear based on the models coming out of China and well see more capabilities in smaller, more efficient models.

Re: GLM-5.2 – How to Run Locally

#105
post #58

Earlier quoted context omitted.

which mac is smoking the spark?

pretty much any of them, dude, as long as you have enough RAM, since it uses unified RAM and a powerful SoC CPU/GPU. Literally any M-class model, but the M5 is currently top tier.

The DGX Spark has basically the same memory bandwidth as a M5 Pro, and far more than a M5.

Only the M3 Ultra really beats it, and once you start scoping out the cost of a M3 Ultra with 128GB or 256GB, the DGX Spark doesn’t look bad after all.

Re: GLM-5.2 – How to Run Locally

#106
post #34

Earlier quoted context omitted.

A nvidia spark thingie has 128GB unified RAM. They also have a dual port version of one of these things: https://www.nvidia.com/content/dam/en-zz/Solutions/networkin... . ie 2 x 100GB/s ports, they may even be 2 x 200GB/s. Once I've got my paws on one, I'll know more. You can cluster these beasts too. Two and three (with two IP subnets) is fairly obvious. Four or more might need a switch depending on how much network…

128 gb of much slower ram than Apple.

DGX Spark is ~273GB/s. That’s about M5 Pro territory, and twice as fast as the M5. You’d have to go to the M5 Max, or M3 Ultra, to get higher memory bandwidth than the Spark.

Re: GLM-5.2 – How to Run Locally

#107
post #7

I feel like the gap is closing to be able to run good enough models locally even for coding and I would assume it could make some companies a bit nervous. Am I wrong about that?

If we didn't have a RAM/GPU shortage right now they would be more nervous than they are. But as it is very few people are going to be able to afford a rig that can run this model effectively. That's probably not going to change for several more years yet. I think if the Z.ai folks decide to come out with a flash version of GLM-5.2 specialized for coding that came in about about 80B params, then the US frontier labs w…

When a large open weight model is released, a lab, startup, or a rich hoist can easily do logit-level distillation and create a XXb param model or whatever, and in theory you should get a really good distill.

Re: GLM-5.2 – How to Run Locally

#108
There is a push from multiple directions at the same time:

- new AI desktops with GB10s. They are relatively cheap and you can cluster them and load 1TB of VRAM

- Nvidia, amd, intel, Cerebras etc pushing new hardware

- oss models getting crazy good, like glm 5.2

- flash models getting very good like deepseek V4 flash

- quantizations

- harnesses being able to use different models (big for difficult stuff, small for grunt work)

So hopefully soon for the ones who want to break free from APIs, we will be able to host at home a cluster of AI desktops at a reasonable price with Opus-level capabilities, can't wait!!

Re: GLM-5.2 – How to Run Locally

#109

Earlier quoted context omitted.

If the newer models require more/better hardware then you’ll lose capabilities. I think you’re better off renting GPU instances and running all the software on those. It’ll be cheaper than Anthropic and OpenRouter but slightly more expensive than electricity and depreciation of hardware.

The newer models don't require more/better hardware. There's a small army of local llm enthusiasts who are running LLMs using 3090s and H100s because they have lots of memory. Them being old isn't really that big of an issue as the compute power needed is relatively low all things considered. The number of parameters needed for these open weight models has mostly stabilized so the actual memory requirements aren't li…

Correct. The main bottleneck with LLM inference is, and have always been, memory bandwidth.

TPS = active weights in GB / your memory bandwidth.

That’s it for decode. That’s all.

Re: GLM-5.2 – How to Run Locally

#110

Can somebody help me understand the Quantization Analysis? It says "dynamic 4-bit UD-Q4_K_XL and dynamic 5-bit UD-Q5_K_XL are generally lossless" while showing a top-1% token agreement on the chart of 97.5%. Not what I would consider "generally lossless". Is this implying that some post-processing is going to account for the 2.5% loss? Beam search?

Generally 97.5% token agreement is very positive. Like the article explains, the difference isn’t the model thinking the capital of France isn’t Paris, but rather maybe saying “The capital of France is Paris” instead of “Paris is the capital of France”.
Post reply on HN