Live data from Hacker News

My local model setup on an M4 Pro Mac Mini

lws.io

181–190 of 207 posts

Re: My local model setup on an M4 Pro Mac Mini

#181

Earlier quoted context omitted.

> Pretty expensive is an understatement. [...] If you could it would be multiple hundreds of thousands of dollars. Obviously, I quantified both the operating expense and the capital expense in my post. What I find curious is that you're quoting me talking about the operating expenditure, and changing the topic to be about the buy-in like these are interchangeable things. You don't think that this is a crucial and imp…

> You could have spent all of 5 seconds of searching rather than just assuming[1]. I guarantee this will not ship to you any time soon. The current lead time on these GPUs in measured in years. If you didn't place an order for this a long time ago, it's not coming this year. Being able to add it to an online configurator does not mean anything right now. > 12 months of Claude burning $70k a month is $840k Your math i…

> I guarantee this will not ship to you any time soon.

The assumption, the starting point, is that you have a line on the hardware. Asking around, some distributors have a 6 month lead time on Instinct GPUs, which curiously enough is about how long you'll be twiddling your thumbs waiting for the cooling loop to be put in. Yes things take time.

> Your math is completely useless with these arbitrary numbers pulled out of the air.

Your dismissal is worthless if you can't even be bothered to provide a counter-example. You've not provided a single iota of quantified reasoning beyond my original not accounting for the space used for the context of concurrent users.

> If you want to begin calculating payback period you'd need to look at token costs, cost per task, utilization rates, and so on.

Now go back and carefully reread my original post. Yes, if you are not actually redlining an LLM for a billing cycle, the capex starts to be way more relevant for this setup. Otherwise, our constraint is time and our unit of measure is $/hr.

If you want to compare token cost, it may shock you to learn that Kimi K3 without speculative decode on this setup is slightly under twice as fast as Opus 4.8 max. That's still true when fast is compared with K3 with speculative decode, and now Claude is twice as expensive as a base rate. Oops. We're already burning more money over a period of time, looking at tokens we're screaming even further ahead.

> You're also neglecting the fact that hosted tokens are going down in price at a rapid rate.

Cool. Call me when Opus 4.8 max is $0.50/million. In 4 years you could have bought the 200 acres of land down the road from your building, started a 5MW solar farm subsidiary that you'll expand over time, and as soon as your connect is up, dropped the opex of the cluster down to its maintenance costs. That subsidiary will pay the loan required to spin it up back irrespective of your primary business. When you own your own shit, you can play your own game, stack your cards deep. Have a little bit of business acumen. Fuck what The Valley is doing, that is an ecosystem fully enslaved by economic nihilism, money isn't grounded there.

> Not nasty or defensive, just tired of these armchair claims that it's easy to go out and buy an 8 X MI355X box from people who obviously have no idea what the hardware lead time is like right now

This motte-bailey routine is both nasty and defensive, particularly when you keep prosecuting a geist of numeric justification that never arrives. All I've gotten from you is vague dismissals, one borderline irrelevant technical argument, moving goalposts and missing the point. Granted, not as egregiously as other people in this chain thinking we're talking about running 100B models on a Mac, I'll give you credit for that. But this whole time, we're just talking past each other. You make realistic points and I try to bring you back to context, but you have to work with me here too.

The point was that these companies are not selling to you below cost, they're not even selling to you at-cost. Just use your head. Venture capital isn't a magic wand. Frontier companies are in the red because they're in non-stop expansion operations at massive scales. Anthropic has an operating profit of half a billion dollars[1]. They are not selling you API usage below cost.

[1] - https://www.forbes.com/sites/jonmarkman/2026/08/17/anthropic...

Re: My local model setup on an M4 Pro Mac Mini

#182
post #3

Earlier quoted context omitted.

I have an M4 pro (48 GB ram) and I run Gemma 4 26b a4b at 52 tok/s and Qwen 3.5b a3b at 72 tok/s. Both 4bit quantized. These are enough for my needs and the performance is more than good enough. I'm not running the MLX version of the Gemma model, if I did the inference speed would likely be a bit better. I wouldn't use them for coding features though.

I have a (now discontinued) 64gb mini pro and I’ve found the same qwen model to be almost unusable unless I kill Thinking on each turn. What are you using them with/for?

I do turn thinking off most of the time for both models. I made a separate comment detailing my use cases.

Re: My local model setup on an M4 Pro Mac Mini

#183
post #162

If author runs the local model for privacy reasons, then I don't understand why they give Telegram access to all their conversations. It's well known that Telegram doesn't end-to-end encrypt bot accounts.

The beauty of this is that you can just swap out the platform and everything remains as it's the same backend. You make a really good point, one that I haven't really considered, but I also only have so many hours in the day to be fiddling around with stuff.

I think I'd use Discord instead. Clankers are happy to set it all up for you.

Re: My local model setup on an M4 Pro Mac Mini

#184
post #124

Earlier quoted context omitted.

It’s not so clear after 5 years that you’ll come out ahead. You’ll have spent $20k. The apple computer owner will probably be running local models that are better than today’s frontier on the same hardware. Idk where you live, but where I am running the M5 Ultra Mac Studio at max rated power 24/7 for a month costs C$42. The considerations against Apple hardware are 1) hardware advancements 2) early access to the best…

> The apple computer owner will probably be running local models that are better than today’s frontier on the same hardware. Hardware is not magically getting more memory or bandwidth. Believing there will be some magical optimizations to compensate for it is just dellusion.

Open weight models have been getting better/smaller every year.

Also, from what I can tell, MLX inference is not as well optimized as CUDA, and the M5 Ultra has additional kinds of AI compute which is unavailable on other M models. With the massive 1.2 TB/s 512GB Mac studios coming out, I think MLX will get a lot more attention.

In short: Todays models should run faster next year, and next year's models should also be more efficient.

Re: My local model setup on an M4 Pro Mac Mini

#185
post #57
post #44

Earlier quoted context omitted.

>You couldn’t buy one of these if you wanted to right now. You can: https://www.exxactcorp.com/Exxact-TS4-149591758-E149591758 . You can get thousands of tps of GLM 5.3 output out of this thing, which grades around Opus 4.8. Payoff is around 1 year vs. spot prices on these GPUs, including power.

I have quoted large nodes from this supplier and have lots of^W^W GPUs from them for personal use. Current lead time is more than 30 months. They're a good provider but you have to be a big shot buying NVL72s before you're getting anything within your payback period.

Yeah, you're looking at the scale of ... Nscale ... buying in the order of 1,000 NVL72 racks to get yourself deliveries in a timely manner.

Re: My local model setup on an M4 Pro Mac Mini

#186
post #153
post #64

Most people running local models would probably love to run larger models if only they had access to big enough hardware. I'm curious: to those of you running models locally, if there was a way to inference the model of your choice at a reasonable cost by effectively time-sharing a B300 rack through some privacy-protecting intermediary, would you consider that? If there was a "Mullvad of GPU clouds", would that solve…

Chutes, Near AI, Phala and Tinfoil all offer various privacy assurances around inference. Some of the bigger providers also offer "zero data retention". The problem I have with these is that the guarantees aren't strong enough (Phala, Near) or the models are old (Tinfoil). Chutes is mostly pretty good (cryptographic security all the way to the GPU) but I'm not sure it's possible to cryptographically verify the precis…

These are all on my router TrustedRouter, and more providers coming. Tinfoil has some newer ones too like GLM 5.3 now.

Phala isn't verifying all the way down but NEAR is and I know the CEO

Re: My local model setup on an M4 Pro Mac Mini

#187
post #64

Most people running local models would probably love to run larger models if only they had access to big enough hardware. I'm curious: to those of you running models locally, if there was a way to inference the model of your choice at a reasonable cost by effectively time-sharing a B300 rack through some privacy-protecting intermediary, would you consider that? If there was a "Mullvad of GPU clouds", would that solve…

i've got a "Router of GPUs" end-to-end encrypted: TrustedRouter.com.

Re: My local model setup on an M4 Pro Mac Mini

#188

Earlier quoted context omitted.

Waste of time when I can pay $20 a month for sol.

versus $0 with local models. There will always be a reason to run frontier models, but local models are well at levels that assist with stuff that don't need that level of complexity.

You could buy a $150 refurbished 16gb i5 and use that model until openAI does its IPO and has to become sane again.

But I guess a $2000 Mac is probably better if you don't care about cost or quality.

Re: My local model setup on an M4 Pro Mac Mini

#189

Earlier quoted context omitted.

I watched someone at a fortune 20 company get embarrassed for buying a Mac to run a 70B model in 2025. He was a lead engineer, so after he announced it wasn't going to work, everyone pretended it never happened. But we all knew.

Sounds like a really rude workplace. Who cares if he wants to try running things locally?

What was rude? No one bothered him.

Also the entire purpose of them buying it was so the department had a LLM.

I proposed A6000. That ended up working.

Re: My local model setup on an M4 Pro Mac Mini

#190
post #106

Nice setup; but, for simple tasks or questions, AI is currently free? And it will probably stay free, as I don't see Google starting to charge for using AI on its search engine? So costs can't be a motivation for running small models locally? For more complex or important tasks, costs, autonomy and privacy matter, but then so does performance/quality. So I'm not completely convinced it's really worth it; but it's tem…

You are correct. In large part, the cost of something like Gemini on a very basic Google AI plan provides far more utility than local LLMs for coding assistance.

There are 2 main reasons for running local LLMS.

1. Process private data/work with uncensored models.

2. Use a large amount of inference that would quickly blow through rate limits and/or run up API costs.

The thing that is critical for 2 is that a) you have to have a sweetspot between a pretty good model, which means largest parameter counts, and fast enough token generation where you can run agentic loops. The latter is needed because you aren't going go get the "intelligence" of larger models to form shell commands and run tools to figure stuff out, so the only way around that is to have custom agentic loops to force the model into doing what you want, which results in more text processing.

From my testing, Gemma4:31b is basically the only local model that can be relied upon to produce accurate results. Qwen models chase benchmarks, which results in MoE models (thus the A3B in the model, i.e 3 billion parameters are only active during inference). In general, these are good for very specific tasks, but fail to be accurate in considering cross task data, whereas Gemma, being fully active does a much better job. If you only need to do a very specific deterministic task, those models are pretty good.

As an aside though, if your task involves pure text processing (for example take html data, make it into a markdown document), you can also additive train Gemma270M quite easily all on CPU, and on a decent CPU it gets like 50-100 tok/sec, no need for any extra hardware.

The thing with Macs is that while they can run those models and larger models no problem, the tok/sec is very slow. This limits effectively what you can do with the models. On the M4 that the poster mentioned, Gemma:31b will run about 20 tok/sec. That means that when you wants to write a whole code file or process large context, you have to wait for it to do things. Compared to workflow with larger models, where file generation often takes like The only benefit of using Macs is the price for Mini and cheaper studios. However, once you reach the total cost of about 2.5k (note that the M4 statedin the article us about 2k), building a gfx card rig is the way to go. You can get 100 tok/sec on a 3090, and it will feel a lot like the cloud models.

Post reply on HN