Live data from Hacker News

GLM-5.2 – How to Run Locally

unsloth.ai

291–300 of 328 posts

Re: GLM-5.2 – How to Run Locally

#291

Earlier quoted context omitted.

if only there was a magical place where geothermal and hydroelectric is ubiquitous and the weather is cold enough that no one is going to be complaining about free heating.

To be fair, Vancouver is such a magical place in terms of electrical cost, but the cost of living and real estate are otherwise through the roof, with decrepit and nasty (would need $100k in renovations immediately if you're not treating it as a teardown) single family detached homes on the east side of the city selling for 3.2 million.

Yeah there's a reason our datacentres are in Kamloops, cheap housing and a big ass river right next to it. It even gets decently cold in the winter so you can save on cooling.

There's also tons of opportunity to build them out in former pulp mill towns on Vancouver Island that have big interconnects or dedicated generation.

You'd have to be an idiot to put a datacentre in Vancouver, or have fuck-off scale monopoly money, which is probably why Telus is doing it.

Re: GLM-5.2 – How to Run Locally

#292
post #262

Earlier quoted context omitted.

if only there was a magical place where geothermal and hydroelectric is ubiquitous and the weather is cold enough that no one is going to be complaining about free heating.

The largest geothermal plant in the world is only 1.5GW, in the United States, which is over double all the plants combined in Iceland. The second largest is 1/3 that, in Mexico. [1] There is no "ubiquitous" geothermal where there also high power usage. Data centers have to go where power is, not can be. [1] https://en.wikipedia.org/wiki/List_of_geothermal_power_stati...

Related, it should surprise no-one that the tech giants are interested in nuclear [1], including small reactors [2], rather than waiting for the utility monopolies [3] to raise an arm and actually generate more power [4].

[1] https://www.cnbc.com/2025/03/12/amazon-google-and-meta-suppo...

[2] https://www.sciencenews.org/article/small-modular-nuclear-re...

[3] https://floodlightnews.org/fraud-and-corruption-on-rise-at-u...

[4] https://decarbonization.visualcapitalist.com/animated-70-yea...

Re: GLM-5.2 – How to Run Locally

#293
post #122

Earlier quoted context omitted.

How noisy does his fan get…

it doesn’t get noisy at all

In case anyone was wondering my spark is basically silent as well. It's great at being ignored, if that's really important to you. I've run mine completely headless since I bought it, including setup.

Re: GLM-5.2 – How to Run Locally

#294

I run Q4_K_XL. All it takes to run to get about 6tk/sec is 512gb of ram and 2 3090 GPUs with llama.cpp -cmoe. I also have crappy DDR4, 2400mhz, 3200mhz will bring that speed up to about 9tk/sec. I also have ok 32core epyc CPU, a better 64core would bring it up to about 11tk/sec. I did a budget build before the crazy hardware cost and I regret it everyday. Nevertheless, it's fantastic being able to run this model at h…

Running that full load is at least 600 W, so in a day ~14 kWh. At $0.2 a kWH, that would be $2.80/day or $1k a year of op-ex in electricity. Unless you really want privacy or the fuzzy feeling of owning your own, it’s cheaper, more convenient and has much faster tok/s if you pay a hyper scaler. That said, I do like the direction we are heading and look forward to seeing what host your own hardware we get in 2 years.

Isn't that still cheaper than the 100 or 200$ plan that Anthropic wants from you?

Re: GLM-5.2 – How to Run Locally

#295
post #220

Earlier quoted context omitted.

Well, it's about GPU VRAM if you want something competitive with cloud-hosted offerings at the performance levels showing in benchmarks. This is a heavy quant with quality degradation and significantly lower performance. Cloud offerings are 80-200tk/sec versus single digit tk/sec. That said, I'm also surprised it runs at all locally. I do think it'd be painfully slow for anything interactive so you're relying on anot…

I see. So not quite usable apart for specific use cases. Maybe in a few years we'll see new hardware players and better prices.

I think we'll see

- better hardware

- more efficient model runtime algorithms/code

- smarter/more efficient models (same capability with less parameters)

So ideally these will all come together and help.

Re: GLM-5.2 – How to Run Locally

#296
post #270
post #162

I have high respect for unsloth's work, helping millions to get started with local AI, but this post appears kind of download bait. Offloading too many layers to CPU is not going to work at all. I have tried this many times and had to rm -rf on those heavy hf cache folders. Also I doubt 1-bit or 2-bit quants of GLM 5.2, running mostly outside of VRAM can beat Q8_0 of Qwen3.6-27B fully loaded in VRAM - on usefulness.

I run 3bit GLM5.2 and full precision Qwen3.6-27B. GLM is much much closer to frontier models in it's breadth and ability to plan. If you just need to implement Python code from an existing plan Qwen is your choice but it has problem succeeding with more complex tasks that GLM5.2 does not. As I type this my local GLM5.2 is troubleshooting bugs that Qwen would not be able to handle.

Not sure how much of your GLM is offloaded to CPU. I was contending the suggestion of using system RAM + VRAM.

Re: GLM-5.2 – How to Run Locally

#297
post #103

Earlier quoted context omitted.

> The ram/gpu shortage won't last forever though. No disagreement there, but it could easily last another 3 to 5 years which is a long time in tech terms.

Long enough for them to IPO and all the execs to retire. I doubt they care beyond the IPO.

thats pretty scary to me. what will the data centers be used for if people run that stuff offline? maybe new models? but will there even be any demand? I guess we will see

Re: GLM-5.2 – How to Run Locally

#298
post #280

Earlier quoted context omitted.

If we didn't have a RAM/GPU shortage right now they would be more nervous than they are. But as it is very few people are going to be able to afford a rig that can run this model effectively. That's probably not going to change for several more years yet. I think if the Z.ai folks decide to come out with a flash version of GLM-5.2 specialized for coding that came in about about 80B params, then the US frontier labs w…

is it possible that ai companies ordered a bunch of ram just so that models cannot be run locally? they are betting new fabs wont be built before quantum takes hold.

I am quite certain that it is delayed on purpose to maximize the gains, but at some point some company will see the huge demand for local ai and will want to eat the cake (given that it is feasible)

Re: GLM-5.2 – How to Run Locally

#299
post #7

I feel like the gap is closing to be able to run good enough models locally even for coding and I would assume it could make some companies a bit nervous. Am I wrong about that?

You don't even need to run them locally for them to be a threat. Plenty of companies are looking at paying third party companies to host these models and they come in at fractions of the price of the frontier labs.

thats true. also, I watched the glm prices and it didnt take long before the prices dropped even lower for some providers. its like another layer of competition between hosters

Re: GLM-5.2 – How to Run Locally

#300
post #34

Earlier quoted context omitted.

A nvidia spark thingie has 128GB unified RAM. They also have a dual port version of one of these things: https://www.nvidia.com/content/dam/en-zz/Solutions/networkin... . ie 2 x 100GB/s ports, they may even be 2 x 200GB/s. Once I've got my paws on one, I'll know more. You can cluster these beasts too. Two and three (with two IP subnets) is fairly obvious. Four or more might need a switch depending on how much network…

I have one, and I love it. That said my buddies Mac smokes it for inference workloads in terms of tokens per second AND its more usable for other things. If you are training and doing research it's great, if you want to cluster them it cant be beat, but if you just want local inference on a single box buy a mac or even a strix halo device.

Get your buddy with his smoking Mac to allow multiple concurrent connections and see how it gets on compared to your Spark. Don't ever use a single "chat" test to derive performance - try running say 10 or more.

You might also notice that your Spark has a pair of QSFP28 or DD (not sure yet) type interfaces as well as the 10Gb/s ethernet - that network card is a right old beast and adds quite a lot to the cost. It is capable of either 200 or 400Mb/s and can be split into two lots of four. Your mate's Mac probably has a wifi connection and is too cool for ethernet 8)

That NIC is there for a good reason - the Spark wants some friends to cluster with and you will absolutely spank any Mac when you spaff Mac style money on say three of these beasts and some cables and cluster them up. If you want four or more, you will need a switch and Mikrotik and others have them.

Casual "tokens per second" in AI is a bit like gamers whittering on about "ping" when they are using TCP and UDP for their games. ICMP request/response is a handy way of testing network paths and can give some indications towards potential performance limitations.

Post reply on HN