Live data from Hacker News

GLM-5.2 – How to Run Locally

unsloth.ai

191–200 of 328 posts

Re: GLM-5.2 – How to Run Locally

#191

Earlier quoted context omitted.

Where I live prices are often higher than 20c/kWh, but lets take your example and halve it (10c/kWh) so it's ~$1.40/day or ~$500/year. On Openrouter, the cheapest GLM 5.2 provider costs $3/MTok (at 44 tps). Assuming most use is output tokens, that's still the equivalent of 450k token/day, so we're in the same ball park, but without the capex for 2 3090's and the machine. Self hosted only makes economic sense if your…

That's true, there's a lot of places where power is considerably more expensive than $0.20 USD/kWh. But also the 600W figure assumes that it's fully loaded 24x7x365. Running a system that will be 600W under max CPU usage on all cores and RAM and a few 3090-class GPUs, that same system might be only 90W or around there when idle at 0.00 unix load. If we say: (600 * 24 * 31)/1000 = 446kWh in a month at full load 24 hou…

> person is using it in bursts and intermittently throughout an 8 hour workday.

You can’t do that with 6 tps, though.

Re: GLM-5.2 – How to Run Locally

#192
post #45

Earlier quoted context omitted.

Funny I casually asked Gemini and it said 500k for unquantized with decent throughput.

i asked gemini and it replied with "Error: 400 Your prompt was blocked by safety filters. Please revise and try again."

Safety from competition!

Re: GLM-5.2 – How to Run Locally

#193
post #134

Earlier quoted context omitted.

But what is it doing for you that you couldn’t do yourself at that speed? I‘m really curious and on the fence of partly going local.

Not a Local LLM user, but I regularly kick off meaty jobs in Claude Code then check on them 1-2hrs later.

In this case it would be 20-40 hours to accomplish the same amount in f work when running locally

Re: GLM-5.2 – How to Run Locally

#194

There is a push from multiple directions at the same time: - new AI desktops with GB10s. They are relatively cheap and you can cluster them and load 1TB of VRAM - Nvidia, amd, intel, Cerebras etc pushing new hardware - oss models getting crazy good, like glm 5.2 - flash models getting very good like deepseek V4 flash - quantizations - harnesses being able to use different models (big for difficult stuff, small for gr…

Hope you're right! Can't wait!

Re: GLM-5.2 – How to Run Locally

#195
post #103

Earlier quoted context omitted.

> The ram/gpu shortage won't last forever though. No disagreement there, but it could easily last another 3 to 5 years which is a long time in tech terms.

Long enough for them to IPO and all the execs to retire. I doubt they care beyond the IPO.

I think this is the play

Re: GLM-5.2 – How to Run Locally

#196

Earlier quoted context omitted.

I don't think so. I could easily see a company deciding to host and run these models for their own development. If you have a dev team of about 10 people, a one time $50k investment in an LLM server has to be pretty tempting. Unlimited tokens, decent performance, upgrade options, and potential product integrations. For companies wanting LLMs in their products in general, I have to think going the local llm route is e…

Surely for most the desire is just an LLM provider that doesnt store or sell their queries (including by national actors). As long as that is allowed to happen surely its the answer for the vast majority.

> LLM provider that doesnt store or sell their queries

> As long as that is allowed to happen

It won't be. Only we can provide that, and only for ourselves.

Re: GLM-5.2 – How to Run Locally

#197

I run Q4_K_XL. All it takes to run to get about 6tk/sec is 512gb of ram and 2 3090 GPUs with llama.cpp -cmoe. I also have crappy DDR4, 2400mhz, 3200mhz will bring that speed up to about 9tk/sec. I also have ok 32core epyc CPU, a better 64core would bring it up to about 11tk/sec. I did a budget build before the crazy hardware cost and I regret it everyday. Nevertheless, it's fantastic being able to run this model at h…

Running that full load is at least 600 W, so in a day ~14 kWh. At $0.2 a kWH, that would be $2.80/day or $1k a year of op-ex in electricity. Unless you really want privacy or the fuzzy feeling of owning your own, it’s cheaper, more convenient and has much faster tok/s if you pay a hyper scaler. That said, I do like the direction we are heading and look forward to seeing what host your own hardware we get in 2 years.

which hyper scaler would you suggest ?

Re: GLM-5.2 – How to Run Locally

#198

I run Q4_K_XL. All it takes to run to get about 6tk/sec is 512gb of ram and 2 3090 GPUs with llama.cpp -cmoe. I also have crappy DDR4, 2400mhz, 3200mhz will bring that speed up to about 9tk/sec. I also have ok 32core epyc CPU, a better 64core would bring it up to about 11tk/sec. I did a budget build before the crazy hardware cost and I regret it everyday. Nevertheless, it's fantastic being able to run this model at h…

LOL, sure this works if one has a time machine or a LOT of money to burn.

32 CPU Epyc (Epyc is required for faster memory access) + 32 GB VRAM + 512 GB RAM is stupid expensive nowadays, and in best case, it will just downgrade to "very" expensive at some point in the future.

This makes sense only if 1. one is paranoid about privacy or 2. they have money to smoke or 3. they need to workaround cloud model restrictions, AND they have to do it routinely (because if not, a oneshot cloud bare metal setup is way cheaper, faster, and allows more powerful models, due to VRAM offering).

I did spend stupid money as well and yet, the system is 2x slower than cloud providers for comparable performance on vision tasks (I still have to test coding). Oh, and it's hot as hell.

Re: GLM-5.2 – How to Run Locally

#199
post #2

So close! My machine with 192GB RAM + RTX 3090 24GB can almost run this. It says it needs 24GB of VRAM and 256GB of RAM for MoE offloading. https://unsloth.ai/docs/models/glm-5.2#usage-guide In a prior thread, someone said it would take $500k in hardware: https://news.ycombinator.com/item?id=48629970

$500k is a vast overestimation. For massive concurrency at FP8 or even BF16 maybe. NVFP4 at reasonable speeds (~120 tok/s) and concurrency is possible at a $80/90k figure with today's prices, maybe even less. That buys you 6 RTX 6000 PRO Blackwells, a decent CPU and motherboard, power supply. 576gb of VRAM. You could do it for under $50k if you're OK with 40 tok/s decode, ~1200 tok/s prefill.

You can get a 1TB of HBM2 vram for like 10k, https://www.ebay.com/itm/177571378959

The problem is the backplane I have not managed to find a single baseboard, and getting a random baseboard to work with random modules is probably a crap shoot.

Re: GLM-5.2 – How to Run Locally

#200
post #165
post #161

Earlier quoted context omitted.

I guess you missed recent news. Problem is that cloud LLM might just sliently sabotage your work by downgrading output model with no notice. Or cloud LLM might just refuse to sell to you because it dont like your passport.

So you're buying expensive hardware as insurance for the case that your cloud provider turns against you and you have to switch to another of the twenty offering the same model https://openrouter.ai/z-ai/glm-5.2 or in the worst case buy the same hardware later? How does that make sense?

It’s rationalization for what people want to do anyway.

Like buying a new car today and taking on gas, parking, etc, expenses in case the bus route you’re using goes away at some point in the future. It’s not an economic decision, it’s a desire to have the new car dressed up in what-ifs.

Post reply on HN