Live data from Hacker News

My local model setup on an M4 Pro Mac Mini

lws.io

161–170 of 208 posts

Re: My local model setup on an M4 Pro Mac Mini

#161
post #157
post #144

Earlier quoted context omitted.

true but if you're actually running k8s and similar workloads, chances are it might eat memory that LLM requires. you'll also notice these articles rarely specify their context window in tokens, because it is small, usually 30k to 70k tokens and it gets slower as it fills up.

I actually have a Mac Mini M4 Pro with 48G. I gave the k8s example because this is what I was doing with it. Was because I am back to using Linux as my workstation. My Mac Mini is now a headless server for llama.cpp. So, you are right that for these workloads , I would not be using the Mac Mini for k8s AND llama. Another thing going against using a Mac for Linux containers is that there are no solutions that I know t…

Try smolvm microvms from https://smolmachines.com - among other benefits they only consume host resources if they're actually used.

Re: My local model setup on an M4 Pro Mac Mini

#163

I just use Gemini Pro which comes bundled with a Chromebook; when the 'free' year runs out I buy another, initiate the free years Gemini again and then put the 'as new' Chromebook on eBay and get most of my money back. So Gemini Pro costs me less than £2 a month. And for my needs (investment research) that works well and is a pretty cheap compromise.

You didn't read the article nor understand the purpose.

The way you are using it uses the internet and datacenters. It is costly to the environment.

Running locally is a significant cost savings in comparison.

Re: My local model setup on an M4 Pro Mac Mini

#164
post #106

Nice setup; but, for simple tasks or questions, AI is currently free? And it will probably stay free, as I don't see Google starting to charge for using AI on its search engine? So costs can't be a motivation for running small models locally? For more complex or important tasks, costs, autonomy and privacy matter, but then so does performance/quality. So I'm not completely convinced it's really worth it; but it's tem…

Traditional search is “free” too, but you see ads. If something looks free, then you are the product.

OpenAI has started doing ads, but for Anthropic the free tier is still just a loss leader. It's basically an ad for their pay tiers.

Re: My local model setup on an M4 Pro Mac Mini

#165

Earlier quoted context omitted.

> With an 8x MI355x cluster at full tilt and including cooling, your power draw runs ~17kW. That's what it looks like when it's running full tilt. To be fair, hey that's pretty expensive. Pretty expensive is an understatement. You couldn’t buy one of these if you wanted to right now. If you could it would be multiple hundreds of thousands of dollars. > It does mean 8 multi-trillion parameter models unquantized runnin…

> Pretty expensive is an understatement. [...] If you could it would be multiple hundreds of thousands of dollars. Obviously, I quantified both the operating expense and the capital expense in my post. What I find curious is that you're quoting me talking about the operating expenditure, and changing the topic to be about the buy-in like these are interchangeable things. You don't think that this is a crucial and imp…

"You can of course trot out the point that oh, in 12 months this setup will be extremely outdated! "

But interestingly still extremely valuable on the second hand market.

The capital expense isn't the amount laid out. It's the rental cost of obtaining that capital, less the depreciation on the fixed asset over the period in use.

Going back to the OP, Apple gear is well know for having good resale values, which means the capital outlay isn't anywhere near as much as some people think.

Re: My local model setup on an M4 Pro Mac Mini

#166

Maybe a silly question, but is there a reason/advantage to using mac minis over any other kind of small computer/laptop, ie running linux? My understanding of using a mac mini for ai (ie running a claw bot or whatever) is to have it 'always on' and a better price/performance profile than a cheap vps. Is there performance (silicon processor?) so unique? As I guess it's not their graphics units. I see tonnes of people…

There are probably three main reasons: 1) unified memory (but you can also that with DGS Spark / Strix Halo), 2) access to your Apple account, so you can have a bot handle your iMessages, email, calendars 3) energy efficiency.

Apart from that, if it doesn't work out you still have a Mac Mini, which in itself is more desirable for many than a DGX Spark or Strix Halo if you have no AI use case.

Re: My local model setup on an M4 Pro Mac Mini

#167

I like the in-depth description. Everything from the naming convention of the models (and how much RAM they require) as well as all the components needed underscores just how complicated this all still is. I suppose I am waiting for AI-in-a-Box to come along so I can (painlessly) join in. (I'm sure wrangling with all these esoteric aspects of LLMs though is fun for some people.)

Hehe I really am just working it out as I go along - I promise it is fairly painless. Hugging Face allow you to specify your machine and then browse models that fit. And then you can just vibe out 'oh this one is a bit slow let me try another' etc

Re: My local model setup on an M4 Pro Mac Mini

#168
post #162

If author runs the local model for privacy reasons, then I don't understand why they give Telegram access to all their conversations. It's well known that Telegram doesn't end-to-end encrypt bot accounts.

The beauty of this is that you can just swap out the platform and everything remains as it's the same backend. You make a really good point, one that I haven't really considered, but I also only have so many hours in the day to be fiddling around with stuff.

Re: My local model setup on an M4 Pro Mac Mini

#170

Maybe a silly question, but is there a reason/advantage to using mac minis over any other kind of small computer/laptop, ie running linux? My understanding of using a mac mini for ai (ie running a claw bot or whatever) is to have it 'always on' and a better price/performance profile than a cheap vps. Is there performance (silicon processor?) so unique? As I guess it's not their graphics units. I see tonnes of people…

No, it's just hype and people cargo culting local LLM guys buying maxed out Mac Studio for its massive and relatively fast GPU-assignable RAM.
Post reply on HN