I’m curious to know if these local AI setups are legitimately useful compared to cloud. I’ve struggled a lot to get something useful out of the hardware I have. I realize I’m somewhat limited (16GB RX 9070), but still, it seems really far off from the kind of experience even a basic $20/month subscription gets me. Any tips anyone might have are appreciated! I’d love to be local first and would be willing to buy hardw…
Apple caught off guard by AI demand for Mac Mini and Mac Studio
271–280 of 636 posts
Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio
#272Earlier quoted context omitted.
32GB of fast unified memory is enough for Qwen 3.8 27B. - 16GB for the weights at Q4 - 9GB for the full 256K context at Q8 - 7GB spare for overhead and system. The problem is that these Macs have 32GB of slow unified memory. Edit: I'm thinking of a headless Mac mini, if you meant running it on the same machine you're using of course you'll need more memory, but LLMs are best served from a headless server so that's wh…
Is this for setup for agentic coding? Why not also run the IDE compiler etc... on the same machine to use those CPU cores as well?
Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio
#273Earlier quoted context omitted.
A simple example. I have an older M2 Mac mini that does the OCR and visual description of all my screenshots. Screenshots are stored on my NAS. I like to screenshot things as a quick way to remember. They are things that I would not be comfortable sending a cloud provider (customer data, prototype screenshots, bank dispute details). It runs Qwen3.5:9b and glm5.2-ocr with Ollama and uses about 10GB of RAM. It automati…
A new base model mac mini is $900. That is 45 month of Gemini. Gemini 4.7 Flash will give better OCR results that Qwen or GLM w/ 10GB.
A $20/month Gemini subscription is truly all you need, then yeah, sure.... obviously a homelab setup is a ridiculous alternative on a pure cost basis. For most people doing "real" work with LLMs 40+ hours per week, a more apt comparison would be one or multiple $200/month subscriptions. At which point the break-even point of a homelab is much sooner.
However, most people running homelabs are doing it for other reasons. Independence, learning, and/or privacy issues.
Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio
#274There is a lot of "AI demand" that isn't just running inference on an LLM whose weights you downloaded. I'm training a model using reinforcement learning with self-play. I can and do use vast.ai when scaling but for experiments it's far faster, and cheaper, to run it locally until the bugs are all figured out. Just provisioning a new instance and copying the relevant checkpoints and things can take 25 minutes. It's z…
Do you find that CoreML manages to fill up your drive with so many tiny files that a reboot takes hours to clean them up? I keep meaning to get my friends still inside the spaceship to file a radar about that.
What game are you building?
Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio
#275There is a lot of "AI demand" that isn't just running inference on an LLM whose weights you downloaded. I'm training a model using reinforcement learning with self-play. I can and do use vast.ai when scaling but for experiments it's far faster, and cheaper, to run it locally until the bugs are all figured out. Just provisioning a new instance and copying the relevant checkpoints and things can take 25 minutes. It's z…
Likewise. I have a huge demand personally to run AI noise-filtering models on many TB per month of raw video files. It takes about 3 days per file. Apples ProRes codec is only licensed to run in high quality mode on a Mac, and so my Nvidia PC can’t do what I need. Thus, I own the beefiest Mac Studio you can currently buy. I would pay more for more TFlops. I have done local LLM on there but it wasn’t interesting. Far…
Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio
#276Earlier quoted context omitted.
> an organization that was primed to recognize and pounce on the next big thing e.g.: previous crypto hype-cycle https://www.pcgamer.com/nvidia-cmp-graphics-card-availabilit...
Crypto was a stupid fad, but Nvidia certainly made a lot of money, so being a vendor to a fad is not stupid.
Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio
#277Earlier quoted context omitted.
Serious question: why not run DGX Spark or Framework Desktop, at 30%-50% lower cost?
M5 Ultra has 4-5x the memory bandwidth of both. 1.2 TB/s memory bandwidth opens up good performance on relatively large models.
Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio
#278I would wait till the ram crisis is over to fetch a future 64gb ram gpu to run Q8 models. Cloud inference until than.
Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio
#279It's fun to see that even an extremely large company can find unexpected product market fit [0]. Per this article, "The company reportedly did not possess an engineering team dedicated to business customers or staff focused on developer relations, and lacked an enterprise AI strategy." That sounds insane in retrospect, but I think there's just inherent uncertainty in what people actually need and will use things for.…
No ‘staff focused on developer relations’ is entirely unsurprising based on what I see from the outside.
“Not as fully staffed as some people might hope” or “Developer Relations isn’t as responsive as I’d like” are both at least not obviously false.
Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio
#280Earlier quoted context omitted.
This is such a tired argument and it seems to be parroted every single time someone talks about local models on hacker news. Yes, of course the most economical path is to hand over all your data and become fully dependent on a cloud provider who is already operating as scale, hoping that they won't change/remove models, hamstring capabilities, or raise prices. If this were a thread about hosting your own email or blo…
It's quite shocking to me how many experienced, tech-savvy people, who used to care about cookies and ad tracking - are now willingly sending their business strategies, highly confidential contracts, and intimate personal issues to a cloud provider because "it is only $0.0x per million tokens!".