Live data from Hacker News

My local model setup on an M4 Pro Mac Mini

lws.io

51–60 of 207 posts

Re: My local model setup on an M4 Pro Mac Mini

#51
post #44

Earlier quoted context omitted.

> With an 8x MI355x cluster at full tilt and including cooling, your power draw runs ~17kW. That's what it looks like when it's running full tilt. To be fair, hey that's pretty expensive. Pretty expensive is an understatement. You couldn’t buy one of these if you wanted to right now. If you could it would be multiple hundreds of thousands of dollars. > It does mean 8 multi-trillion parameter models unquantized runnin…

>You couldn’t buy one of these if you wanted to right now. You can: https://www.exxactcorp.com/Exxact-TS4-149591758-E149591758 . You can get thousands of tps of GLM 5.3 output out of this thing, which grades around Opus 4.8. Payoff is around 1 year vs. spot prices on these GPUs, including power.

> You can: https://www.exxactcorp.com/Exxact-TS4-149591758-E149591758 .

No, you can get a quote for possibly being allocated one in the distant future.

The backlog for these is huge. You cannot buy one any time soon.

Re: My local model setup on an M4 Pro Mac Mini

#52
post #44

Earlier quoted context omitted.

>You couldn’t buy one of these if you wanted to right now. You can: https://www.exxactcorp.com/Exxact-TS4-149591758-E149591758 . You can get thousands of tps of GLM 5.3 output out of this thing, which grades around Opus 4.8. Payoff is around 1 year vs. spot prices on these GPUs, including power.

I can't tell from the ad -- it says "supports" 8x MI350X GPUs, but does that mean "includes" 8x MI350X GPUs? For $300K I'd certainly hope so, but I'm assuming not. A system with 4x RTX 6000s costs about $60K these days, and can (as you note) trade blows with Opus 4.8 if not Fable. In fact, it'll give you a better pelican than Fable 5.1, and in less time.

> trade blows with Opus 4.8 if not Fable.

Okay I love the open models, but the hype is getting ridiculous. The models you can run on 4 X RTX6000 are not Fable level.

Re: My local model setup on an M4 Pro Mac Mini

#53
post #3
post #2

No mention of the performance of the models? I'm able to load a bunch of different models on my little mini-PC with 16GB RAM, but the performance is terrible. I always wonder what performance people are getting with local models that they find is acceptable?

I have an M4 pro (48 GB ram) and I run Gemma 4 26b a4b at 52 tok/s and Qwen 3.5b a3b at 72 tok/s. Both 4bit quantized. These are enough for my needs and the performance is more than good enough. I'm not running the MLX version of the Gemma model, if I did the inference speed would likely be a bit better. I wouldn't use them for coding features though.

> enough for my needs

Which are...?

Re: My local model setup on an M4 Pro Mac Mini

#54

Earlier quoted context omitted.

That's not even remotely close to being true, even once you account for capex. You have to look at the actual usage, look at the token limits. Even if you're paying Anthropic $200k/month for scale-tier, you're going to blow through your token limits trying to run max output 24/7. Three users running Opus 4.8 at max non-stop will probably clean your monthly allowance from daddy Dario in less than a week. With an 8x MI…

> With an 8x MI355x cluster at full tilt and including cooling, your power draw runs ~17kW. That's what it looks like when it's running full tilt. To be fair, hey that's pretty expensive. Pretty expensive is an understatement. You couldn’t buy one of these if you wanted to right now. If you could it would be multiple hundreds of thousands of dollars. > It does mean 8 multi-trillion parameter models unquantized runnin…

> Pretty expensive is an understatement. [...] If you could it would be multiple hundreds of thousands of dollars.

Obviously, I quantified both the operating expense and the capital expense in my post. What I find curious is that you're quoting me talking about the operating expenditure, and changing the topic to be about the buy-in like these are interchangeable things. You don't think that this is a crucial and important distinction?

> You couldn’t buy one of these if you wanted to right now.

You could have spent all of 5 seconds of searching rather than just assuming[1]. You're not buying an Nvidia Superpod™.

> You can’t even run one unquantized multi-trillion parameter (>=2T) model on 8 x MI355x with enough context for concurrent users.

That's certainly fair a point. Although in the English language, especially in legal contexts, the multi- prefix is used inclusively for fractional values. That is it's strictly >1, not >=2. IE an 18 month contract is a multi-year contract, or a $1.6 million dollar asset is a "multi-million" dollar asset. But this is uninteresting semantics.

You are right, but it also doesn't matter. The gap is just that big. You can run 1 single user of Kimi K3 and still not even come remotely close to the $70k or so that a single Opus 4.8 user can burn over the course of a month on left on max. An honestly lowballed amount I know from anecdote. The per-token cost is just really expensive.

> Your math is way off across this post.

You made one technical point above, one that doesn't ever arrive at a relevant rebuttal to the substance of my post. But please, I'd love to hear you elaborate, especially because I didn't actually give much math at all.

If you want math though, here's the math. Let's say you are paying a ridiculous amount of money for electricity, a price nobody in the US pays -- $2 per kilowatt hour. That's about 5x the average rate in California, 4x as in Hawai'i. 17kW @ $2/kWh * ~8766 hours in a year puts that cluster's electrical costs at just shy of ~$298k annually assuming it takes no breaks. Let's make matters worse and round that up to $300k. It's also assuming you didn't invest in a solar hookup for your building, which I don't know why you haven't at this point, especially if you're installing a CDU for your new cluster. 12 months of Claude burning $70k a month is $840k. For a buy in of, you know what, let's call it $500k. Why not? It still doesn't matter. The operating cost is so much lower it's paid for itself plus an additional $40k in the first year. Even at a ridiculous penalty in electricity that nobody pays, even overinflating the amount of money you'd pay for the cluster and the infrastructure to get it set up, it's not even remotely close for a single user where the gap is smaller (IE, you're not wasting "a million dollars" in a year by maxing out the $200k scaling limit every month)

You can of course trot out the point that oh, in 12 months this setup will be extremely outdated! It doesn't matter. If the work it was doing today was useful, it will be useful next year too. And with the rapidly encroaching diminishing returns from parameter scaling, you're probably going to be just fine for a while. Maybe grab a quantized version of a newer Chinese model at the end, before grabbing a newer generation of AMD node. Those MI400s are looking pretty sweet after all.

> If replacing an Anthropic subscription for a whole company was as easy as buying a box for the office and then breaking even in 2 months

If you're locked in, then you're locked in. But don't pretend like you're saving money. You're not.

> it wouldn’t be some little secret that we only discover in a comment online.

Why does this have you so nasty and defensive? It's not a "little secret" that running your own infrastructure is cheaper. Of course it is. You know what else is cheaper? Owning your own office building out in the sticks, rather than leasing part of one in the city. Not everybody can make that work, there are no free lunches after all.

History repeats, these same exact lines were rolled out ad nauseum during the cloud craze. Datacenters are businesses, not charities. Frontier companies rent quite a fair amount of their infrastructure. Even if they resold that compute below cost (they don't), there's a pretty steep cliff before the economics start to look attractive.

[1] - https://www.avadirect.com/GIGABYTE-G893-ZX1-AAX4-Dual-AMD-EP...

Re: My local model setup on an M4 Pro Mac Mini

#55

Earlier quoted context omitted.

I can't tell from the ad -- it says "supports" 8x MI350X GPUs, but does that mean "includes" 8x MI350X GPUs? For $300K I'd certainly hope so, but I'm assuming not. A system with 4x RTX 6000s costs about $60K these days, and can (as you note) trade blows with Opus 4.8 if not Fable. In fact, it'll give you a better pelican than Fable 5.1, and in less time.

> trade blows with Opus 4.8 if not Fable. Okay I love the open models, but the hype is getting ridiculous. The models you can run on 4 X RTX6000 are not Fable level.

Well, they are if you're into animating pelicans. :-P But yes, in the general case Opus is a better match.

And Opus is no slouch. I'm satisfied that GLM 5.3 is just as strong as Opus. Z.AI has promised/bragged that they will be at Fable 5.0 level by the end of the year or early next year, and I don't see any reason to doubt them.

Re: My local model setup on an M4 Pro Mac Mini

#56
post #44

Earlier quoted context omitted.

>You couldn’t buy one of these if you wanted to right now. You can: https://www.exxactcorp.com/Exxact-TS4-149591758-E149591758 . You can get thousands of tps of GLM 5.3 output out of this thing, which grades around Opus 4.8. Payoff is around 1 year vs. spot prices on these GPUs, including power.

> You can: https://www.exxactcorp.com/Exxact-TS4-149591758-E149591758 . No, you can get a quote for possibly being allocated one in the distant future. The backlog for these is huge. You cannot buy one any time soon.

Ah gotcha. Have you tried to order something like this in the past?

Re: My local model setup on an M4 Pro Mac Mini

#57
post #44

Earlier quoted context omitted.

> With an 8x MI355x cluster at full tilt and including cooling, your power draw runs ~17kW. That's what it looks like when it's running full tilt. To be fair, hey that's pretty expensive. Pretty expensive is an understatement. You couldn’t buy one of these if you wanted to right now. If you could it would be multiple hundreds of thousands of dollars. > It does mean 8 multi-trillion parameter models unquantized runnin…

>You couldn’t buy one of these if you wanted to right now. You can: https://www.exxactcorp.com/Exxact-TS4-149591758-E149591758 . You can get thousands of tps of GLM 5.3 output out of this thing, which grades around Opus 4.8. Payoff is around 1 year vs. spot prices on these GPUs, including power.

I have quoted large nodes from this supplier and have lots of^W^W GPUs from them for personal use. Current lead time is more than 30 months.

They're a good provider but you have to be a big shot buying NVL72s before you're getting anything within your payback period.

Re: My local model setup on an M4 Pro Mac Mini

#58

Earlier quoted context omitted.

It does make me wonder how the hosted stuff is so cheap. For pretty much everything else, hosted/rented is more expensive but offers better convenience and flexibility. But for AI, even if you consider the total lifetime cost and are utilizing it heavily. You never break even by buying.

They're not cheap at all. I did one xhigh Qwen 3.8 27B agentic coding task last week via OpenRouter and it cost me like $10. 99% of the cost was in input tokens, I only used like 100k ish output tokens. It was a one shot task asking the agent to implement proxy injection to Guice. It did a pretty amazing job. If you were to use hosted LLMs for a lot of agentic coding, a maxed out M5 Ultra Mac Studio would pay for its…

Qwen is weirdly expensive. Deepseek v4 flash is dirt cheap. You'd need at least 128gb of ram to run this model and in my experience, a days work with it costs around 80 cents.

Re: My local model setup on an M4 Pro Mac Mini

#59

Earlier quoted context omitted.

That's not even remotely close to being true, even once you account for capex. You have to look at the actual usage, look at the token limits. Even if you're paying Anthropic $200k/month for scale-tier, you're going to blow through your token limits trying to run max output 24/7. Three users running Opus 4.8 at max non-stop will probably clean your monthly allowance from daddy Dario in less than a week. With an 8x MI…

Here’s an experiment: purchase an anthropic pro max subscription for $200/m. Now go buy the hardware to run DeepSeek’s equivalent. In a year, who spent more?

In normal times in which hardware used to depreciate (lately that's not the case and HW even appreciates, but let's not get distracted), if you calculate only with depreciation costs, plus the fact that when you have such a setup, it'd take many 200$ subs to cover your lack of limits in the other, I think it'd not be a clear victory for any side.

If you just ask "who spent more in the first year" (100% depreciation) then even with 5-6 max accounts, buying HW will be a couple of times more expensive. But when does it make sense to ask that question?

Maybe the SotA models will need better hardware so your investment will not be useful after a year or you'd need very expensive upgrades? But then (as in Fable case) subscribers need to spend more too.

Re: My local model setup on an M4 Pro Mac Mini

#60
post #57
post #44

Earlier quoted context omitted.

>You couldn’t buy one of these if you wanted to right now. You can: https://www.exxactcorp.com/Exxact-TS4-149591758-E149591758 . You can get thousands of tps of GLM 5.3 output out of this thing, which grades around Opus 4.8. Payoff is around 1 year vs. spot prices on these GPUs, including power.

I have quoted large nodes from this supplier and have lots of^W^W GPUs from them for personal use. Current lead time is more than 30 months. They're a good provider but you have to be a big shot buying NVL72s before you're getting anything within your payback period.

Ah thanks for the solid info, too bad. I'd seen them come up as a pretty good price for 6000 RTX's in the past, which seem generally pretty available, good source for those?
Post reply on HN