Earlier quoted context omitted.
true but if you're actually running k8s and similar workloads, chances are it might eat memory that LLM requires. you'll also notice these articles rarely specify their context window in tokens, because it is small, usually 30k to 70k tokens and it gets slower as it fills up.
I actually have a Mac Mini M4 Pro with 48G. I gave the k8s example because this is what I was doing with it. Was because I am back to using Linux as my workstation. My Mac Mini is now a headless server for llama.cpp. So, you are right that for these workloads , I would not be using the Mac Mini for k8s AND llama. Another thing going against using a Mac for Linux containers is that there are no solutions that I know t…
My local model setup on an M4 Pro Mac Mini
161–170 of 207 posts
Re: My local model setup on an M4 Pro Mac Mini
#162Re: My local model setup on an M4 Pro Mac Mini
#163I just use Gemini Pro which comes bundled with a Chromebook; when the 'free' year runs out I buy another, initiate the free years Gemini again and then put the 'as new' Chromebook on eBay and get most of my money back. So Gemini Pro costs me less than £2 a month. And for my needs (investment research) that works well and is a pretty cheap compromise.
The way you are using it uses the internet and datacenters. It is costly to the environment.
Running locally is a significant cost savings in comparison.
Re: My local model setup on an M4 Pro Mac Mini
#164Nice setup; but, for simple tasks or questions, AI is currently free? And it will probably stay free, as I don't see Google starting to charge for using AI on its search engine? So costs can't be a motivation for running small models locally? For more complex or important tasks, costs, autonomy and privacy matter, but then so does performance/quality. So I'm not completely convinced it's really worth it; but it's tem…
Traditional search is “free” too, but you see ads. If something looks free, then you are the product.
Re: My local model setup on an M4 Pro Mac Mini
#165Earlier quoted context omitted.
> With an 8x MI355x cluster at full tilt and including cooling, your power draw runs ~17kW. That's what it looks like when it's running full tilt. To be fair, hey that's pretty expensive. Pretty expensive is an understatement. You couldn’t buy one of these if you wanted to right now. If you could it would be multiple hundreds of thousands of dollars. > It does mean 8 multi-trillion parameter models unquantized runnin…
> Pretty expensive is an understatement. [...] If you could it would be multiple hundreds of thousands of dollars. Obviously, I quantified both the operating expense and the capital expense in my post. What I find curious is that you're quoting me talking about the operating expenditure, and changing the topic to be about the buy-in like these are interchangeable things. You don't think that this is a crucial and imp…
But interestingly still extremely valuable on the second hand market.
The capital expense isn't the amount laid out. It's the rental cost of obtaining that capital, less the depreciation on the fixed asset over the period in use.
Going back to the OP, Apple gear is well know for having good resale values, which means the capital outlay isn't anywhere near as much as some people think.
Re: My local model setup on an M4 Pro Mac Mini
#166Maybe a silly question, but is there a reason/advantage to using mac minis over any other kind of small computer/laptop, ie running linux? My understanding of using a mac mini for ai (ie running a claw bot or whatever) is to have it 'always on' and a better price/performance profile than a cheap vps. Is there performance (silicon processor?) so unique? As I guess it's not their graphics units. I see tonnes of people…
Apart from that, if it doesn't work out you still have a Mac Mini, which in itself is more desirable for many than a DGX Spark or Strix Halo if you have no AI use case.
Re: My local model setup on an M4 Pro Mac Mini
#167I like the in-depth description. Everything from the naming convention of the models (and how much RAM they require) as well as all the components needed underscores just how complicated this all still is. I suppose I am waiting for AI-in-a-Box to come along so I can (painlessly) join in. (I'm sure wrangling with all these esoteric aspects of LLMs though is fun for some people.)
Re: My local model setup on an M4 Pro Mac Mini
#168If author runs the local model for privacy reasons, then I don't understand why they give Telegram access to all their conversations. It's well known that Telegram doesn't end-to-end encrypt bot accounts.
Re: My local model setup on an M4 Pro Mac Mini
#169> oMLX Is that supposed to be hallucination? The human or other kind. Feels like a made up URL. It's .ai, isn't it?
Re: My local model setup on an M4 Pro Mac Mini
#170Maybe a silly question, but is there a reason/advantage to using mac minis over any other kind of small computer/laptop, ie running linux? My understanding of using a mac mini for ai (ie running a claw bot or whatever) is to have it 'always on' and a better price/performance profile than a cheap vps. Is there performance (silicon processor?) so unique? As I guess it's not their graphics units. I see tonnes of people…