Live data from Hacker News

GLM-5.2 – How to Run Locally

unsloth.ai

51–60 of 328 posts

Re: GLM-5.2 – How to Run Locally

#51
post #5
post #2

So close! My machine with 192GB RAM + RTX 3090 24GB can almost run this. It says it needs 24GB of VRAM and 256GB of RAM for MoE offloading. https://unsloth.ai/docs/models/glm-5.2#usage-guide In a prior thread, someone said it would take $500k in hardware: https://news.ycombinator.com/item?id=48629970

I have the RAM, but not the VRAM. What kind of speed/tps could you expect from a 3090 with 24GBs of RAM? I am somewhat tempted to pick a GPU with 24GBs of RAM.

Generation is basically just memory bandwidth math.

Each token has to read all the active weights. I think that's around 40B parameters active. At a 4-bit quant that's 20GB. With 100GB/s (replace with whatever your bandwidth is) and you get 5 tokens per second.

Re: GLM-5.2 – How to Run Locally

#52
post #34

"it can fit" on 256GB of RAM, but it will be heavily quantized and still run very slowly. The headline number is not token generation, its prompt processing. So if you get 10 tok/s and an API gives you 20-30 tok/s, it doesn't seem that bad on its face, but a mac studio or any other machine that's not loading all of it into GPU will do PP 20-50X slower than a purely GPU based setup, which is what actually makes this u…

A nvidia spark thingie has 128GB unified RAM. They also have a dual port version of one of these things: https://www.nvidia.com/content/dam/en-zz/Solutions/networkin... . ie 2 x 100GB/s ports, they may even be 2 x 200GB/s. Once I've got my paws on one, I'll know more. You can cluster these beasts too. Two and three (with two IP subnets) is fairly obvious. Four or more might need a switch depending on how much network…

I have one, and I love it. That said my buddies Mac smokes it for inference workloads in terms of tokens per second AND its more usable for other things.

If you are training and doing research it's great, if you want to cluster them it cant be beat, but if you just want local inference on a single box buy a mac or even a strix halo device.

Re: GLM-5.2 – How to Run Locally

#54

Earlier quoted context omitted.

equilibrium in one or two more years on the consumer/prosumer side think Apple M6 or M7 with a currently unforeseen denser memory style, 256gb RAM a couple inference or cache improvements on the algorithmic side, using less ram for context windows and doubling token speed again denser open source models, packing more experts for smaller active layers it'll still be expensive but like $8,000 - $13,000 instead of $450,…

Fairly certain that model sizes and computational requirements will grow as the price for LLM compute drops.

Maybe there's a conversation to be had about how much is enough... Unless something beyond my imagination happened, I would be happy enough with Opus 4.5 levels of productivity

Re: GLM-5.2 – How to Run Locally

#55
post #23

Earlier quoted context omitted.

> I am very excited for local LLMs I think we may have GPT 5.5-xhigh level of performance for under 2000 EUR We are maybe 10 years off that. RAM prices are going to continue to increase for the next 2 years at least. Even putting that aside it's currently around 40-70,000 EUR to run this with a FP8 quantization (which you need to get close to maximum performance). To actually get GPT 5.5-xhigh performance in the real…

I wonder, if in the near future any acquisitions of some RAM producers with intent to just keep RAM prices up, will happen from the AI companies. It could seriously hurt their business, if companies will be able to host their AI in some time.

I think AI companies have enough things to spend capital on already.

Re: GLM-5.2 – How to Run Locally

#56

Earlier quoted context omitted.

equilibrium in one or two more years on the consumer/prosumer side think Apple M6 or M7 with a currently unforeseen denser memory style, 256gb RAM a couple inference or cache improvements on the algorithmic side, using less ram for context windows and doubling token speed again denser open source models, packing more experts for smaller active layers it'll still be expensive but like $8,000 - $13,000 instead of $450,…

Fairly certain that model sizes and computational requirements will grow as the price for LLM compute drops.

have you seen the open source LLM space? people fulfill all niches and there are active communities at every range of RAM and all are looking for the most capable in their respective range

a lot of innovation occurring

Re: GLM-5.2 – How to Run Locally

#58
post #34

Earlier quoted context omitted.

A nvidia spark thingie has 128GB unified RAM. They also have a dual port version of one of these things: https://www.nvidia.com/content/dam/en-zz/Solutions/networkin... . ie 2 x 100GB/s ports, they may even be 2 x 200GB/s. Once I've got my paws on one, I'll know more. You can cluster these beasts too. Two and three (with two IP subnets) is fairly obvious. Four or more might need a switch depending on how much network…

I have one, and I love it. That said my buddies Mac smokes it for inference workloads in terms of tokens per second AND its more usable for other things. If you are training and doing research it's great, if you want to cluster them it cant be beat, but if you just want local inference on a single box buy a mac or even a strix halo device.

which mac is smoking the spark?

Re: GLM-5.2 – How to Run Locally

#59
post #2

So close! My machine with 192GB RAM + RTX 3090 24GB can almost run this. It says it needs 24GB of VRAM and 256GB of RAM for MoE offloading. https://unsloth.ai/docs/models/glm-5.2#usage-guide In a prior thread, someone said it would take $500k in hardware: https://news.ycombinator.com/item?id=48629970

$500k is a vast overestimation. For massive concurrency at FP8 or even BF16 maybe. NVFP4 at reasonable speeds (~120 tok/s) and concurrency is possible at a $80/90k figure with today's prices, maybe even less. That buys you 6 RTX 6000 PRO Blackwells, a decent CPU and motherboard, power supply. 576gb of VRAM. You could do it for under $50k if you're OK with 40 tok/s decode, ~1200 tok/s prefill.

Yes, a single GB300 workstation also does it, probably even more than 120tok/s.

Official price 85k...

Re: GLM-5.2 – How to Run Locally

#60
post #43

Earlier quoted context omitted.

$500k is a vast overestimation. For massive concurrency at FP8 or even BF16 maybe. NVFP4 at reasonable speeds (~120 tok/s) and concurrency is possible at a $80/90k figure with today's prices, maybe even less. That buys you 6 RTX 6000 PRO Blackwells, a decent CPU and motherboard, power supply. 576gb of VRAM. You could do it for under $50k if you're OK with 40 tok/s decode, ~1200 tok/s prefill.

How fast will the hardware become outdated? Are there big improvements expected in the next 3 years?

M5 Ultra will ship before end of year, likely. Though with current RAM shortage, likely max spec will be 256GB and in short supply.

In late 2027 or early 2028, Nvidia will release Vera Rubin DGX Spark, likely with double or better the performance of current Blackwell, though unclear if memory capacity will go up much from current 128GB. Two to four of those will run models like this decently.

In 2028 we should expect Vera Rubin RTX discrete lineup, including the replacement to the RTX PRO 6000. Likely memory spec will be minimum 128GB. Good chance of up to 200GB. Two to four of those will run NVFP4 models in this class very well.

Post reply on HN