Live data from Hacker News

GLM-5.3 is now open-weight

huggingface.co

261–270 of 296 posts

Re: GLM-5.3 is now open-weight

#261
post #75

Earlier quoted context omitted.

When we consider: * LLM usage is new for the world * Models are evolving quickly with high worldwide competition * Hardware is evolving despite RAM shortages Is investing a huge sum of money in equipment for local inference a wise use of money? Or are M5 Ultra and equivalently priced local inference hardware future-proof enough to be worth it relative to how the market is evolving? Maybe it’s all a question of what y…

It is absolutely not worth buying hardware to run models for purely (long term) cost reasons. For open weights models the economies of scale means the cloud beats local significantly and your payback time is like 10 years. However there are other reasons (e.g. privacy) that might make it worth running locally for some people.

It seems absurdly naive to rely on "oh, the cloud AI of the future will definitely be as open and priced the same way it is right now."

And not "Hey, these companies have a history of giving you something nice now, and rugpulling you either in quality or price later."

Your "absolutely" seems silly.

Re: GLM-5.3 is now open-weight

#262
post #234

Earlier quoted context omitted.

I’ll be very curious what you get with DDR4. I also almost went that way. I have an Epyc DDR 5 rig and the best I see is 10 tok/s. Caveat being that’s at Q8 and a 4090 doing pre fill so it could be pushed up. The surprising thing for me is how much work you will need to cool the banks if you’re near your memory ceiling. My memory starts soft throttling at about 74C (dies may be hotter, that’s the bank temp) and will…

Typically computers with these larger memory amounts have fans that scream like a banshee trying to move impossible amounts of air over the memory and CPU. Getting something both cool and quite can be a bit difficult.

I built a dual epyc server with 64 cores and 1 TB of DDR4. Draws around 800W or so under load. I used off the shelf liquid cooling. It is audible but not noisy.

The trick is to turn on the cooler's RGB in your 6000€ server to get a free speed boost. I am not liable for sysadmin's heart attack upon reading this.

Re: GLM-5.3 is now open-weight

#263

Earlier quoted context omitted.

> There is a certain personality type that cannot tolerate any uncertainty and must lock everything in right now against all future possibilities. I thought HN banned personal attacks. I'm in this sentence and I don't like it. /s I just buy the good apple hardware because it's good, and it also happens to run local models. It's not as good for the dollar, don't get me wrong, but I'm not going to develop iOS without a…

Welcome to production software, where you really want to pin all uncertainties and dependencies, and roll back in case a major problem occurs.

Nobody is treating production software like that today. It's always downloading half the internet on every build.

Re: GLM-5.3 is now open-weight

#264

Earlier quoted context omitted.

When we consider: * LLM usage is new for the world * Models are evolving quickly with high worldwide competition * Hardware is evolving despite RAM shortages Is investing a huge sum of money in equipment for local inference a wise use of money? Or are M5 Ultra and equivalently priced local inference hardware future-proof enough to be worth it relative to how the market is evolving? Maybe it’s all a question of what y…

I have a Strix Halo and dual 32GB GPUs in my desktop, that sit idle right now, because the electricity to run them and to cool them in 110F weather Texas is currently experiencing pretty much nulls any savings I might see over getting better models from cloud providers. While I mostly use Claude or Codex with subscriptions for agentic work, for API use DeepSeek has usually been my go to, but now I guess it's GLM 5.3…

Solar panels are produced for less than $1 per watt btw. Shame they're illegal in Texas because they compete with the governing oil industry.

Re: GLM-5.3 is now open-weight

#265
post #230

Earlier quoted context omitted.

I wish Texas would write up a regulation allowing 'balcony solar' as I could easily generate 1000-2000w of solar in my small back yard to take a bite out the sizeable cooling bill I have.

Seems like it's easier to ask forgiveness than permission. And, I wouldn't bet on this legislature ever doing anything that would disempower fossil energy or reduce their profits, even a little bit.

There are good reasons you aren't allowed to plug random power generators into the grid. You might be allowed to have your own ones not connected to the grid. Remember that graphics cards run off poorly regulated 12V, although you'd want to regulate it anyway because they're expensive to replace if I'm wrong.

Re: GLM-5.3 is now open-weight

#266
post #140

Earlier quoted context omitted.

When we consider: * LLM usage is new for the world * Models are evolving quickly with high worldwide competition * Hardware is evolving despite RAM shortages Is investing a huge sum of money in equipment for local inference a wise use of money? Or are M5 Ultra and equivalently priced local inference hardware future-proof enough to be worth it relative to how the market is evolving? Maybe it’s all a question of what y…

Jalapeno is matching or very near Vera Rubin at 1/4 the power. I would not buy hardware now.

And Tenstorrent greatly exceeds it, but how will you actually get one of those cards?

Re: GLM-5.3 is now open-weight

#267

Earlier quoted context omitted.

> Tools vs services in my mind. There is no guarantee any provider will continue to do what they are doing for you at the price they are doing it. with open models, there is ecosystem/market of providers, where you can easily switch to provider you like

Until there's an executive order that blocks one model from being served.

such order can target any provider (including closed models) as we know.

Re: GLM-5.3 is now open-weight

#268

Earlier quoted context omitted.

I think that's overly pessimistic. Here's [1] a video of somebody running it on a ~$6000 rig and getting around 14T/s for complex prompts (about double that for simpler prompts). Payback time is going to depend on your electric cost/consumption. In most domains cloud providers end up charging a significant premium rather than a offering a scale enabled discount, relative to local at retail costs. That will almost cer…

With roughly 2.7 million seconds per month, times 14 tokens per second, you are getting 38.5 million tokens a month at most. That’s less than 164USD worth of GLM5.3 tokens on the inference market. So that 6000 USD rig will take 3 years to break even - and only if it runs continuously. And this is being generous, as it’s not even taking quantisation into account.

I think the “killer app” is doing inference without sending the data to China or the US. At home it’s overkill but imagine you are an EU consultancy with a lot of client data to work on, or a company/institution with a lot of sensitive data, buying the hardware to make sure the data stays private is a huge benefit. So is that you “own” the model. Its capabilities, price or access don’t change at someone else’s whim.

Re: GLM-5.3 is now open-weight

#269

Earlier quoted context omitted.

Not OP, but I’m running local models on a M1 Max as well with 64GB RAM. It varies by model, but I’m getting 50-60 t/s with Qwen 3.6 35B and Qwen 3 coder 30B. I’ve also used Qwen 3.8 27B but I get 10t/s on it. It’s useable in some use cases, but I rely mostly on my $20 Claude subscription.

That's so cool. I wonder if the regular M5 can run those models too.

I run qwen 3.8 27b on my m5 mbp, with 48gb of unified ram and I’m getting around 10-15 tok/s.

3.6 35b a3b, I’m getting upwards of 100

Re: GLM-5.3 is now open-weight

#270
post #75

Earlier quoted context omitted.

It is absolutely not worth buying hardware to run models for purely (long term) cost reasons. For open weights models the economies of scale means the cloud beats local significantly and your payback time is like 10 years. However there are other reasons (e.g. privacy) that might make it worth running locally for some people.

I think that's overly pessimistic. Here's [1] a video of somebody running it on a ~$6000 rig and getting around 14T/s for complex prompts (about double that for simpler prompts). Payback time is going to depend on your electric cost/consumption. In most domains cloud providers end up charging a significant premium rather than a offering a scale enabled discount, relative to local at retail costs. That will almost cer…

Does it make sense running 1-bit models for agentic tasks?
Post reply on HN