Live data from Hacker News

Apple caught off guard by AI demand for Mac Mini and Mac Studio

macrumors.com

161–170 of 636 posts

Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio

#161

I’m curious to know if these local AI setups are legitimately useful compared to cloud. I’ve struggled a lot to get something useful out of the hardware I have. I realize I’m somewhat limited (16GB RX 9070), but still, it seems really far off from the kind of experience even a basic $20/month subscription gets me. Any tips anyone might have are appreciated! I’d love to be local first and would be willing to buy hardw…

From what I’ve been seeing, the Mac studios do look like they have potential. I was looking to drop $10k-$15k on one until recently. After comparing a Radeon 7900 XTX vs Ryzen Halos 128GB vs M1 MacBook Pro 64Gb, I landed on just getting an external closure setup with Nvidia RTX 5090.

The model I’m specifically targeting to use at high speeds is Qwen 3.8 27b @q4ks. This model actually proved to be good at coding (it sits somewhere between Sonnet 5 and Opus 5 capability). M1 got 10 tok/s, Ryzen Halo 20tok/s, and Radeon 7900 XTX 50tok/s (can only do 128k context window in Radeon card).

The prefill gets extremely slow around 50k tokens in context window (whatever prompt processing stage entails could be wrong about phases here). It takes about 2 hours to fill the context.

Even with a drafter model intended for speed instead of mtp, I can’t get past 70tok/s, still is extremely slow to process prompts as context grows, and drops down to 40-50tok/s anyway making this config still moot for improvement on my Radeon card.

The only thing I can point to slowing me down is bandwidth of the card itself.

I am waiting to actually get my 5090 right now and I am betting that the 1700 Gbps of capacity will fix my prompt processing speeds. I don’t need full PCIe lane bandwidth to serve my house I just need to load the full model into vRAM and let the GPU do its thing.

Additional benefit to the external enclosure route is being able to migrate the inference between devices more easily. I can develop out the infrastructure then migrate the card to be hooked up to a shared node in the house with all the tools necessary for my family to take advantage of the privacy enhancement that comes with local inference.

Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio

#162

I’m curious to know if these local AI setups are legitimately useful compared to cloud. I’ve struggled a lot to get something useful out of the hardware I have. I realize I’m somewhat limited (16GB RX 9070), but still, it seems really far off from the kind of experience even a basic $20/month subscription gets me. Any tips anyone might have are appreciated! I’d love to be local first and would be willing to buy hardw…

Also, the $20/month subscriptions are HEAVILY subsidized, so it's not an apples-to-apples comparison really

It is a completely reasonable comparison for me as a consumer, since they're the costs and benefits that I'll actually get.

Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio

#163

I’m curious to know if these local AI setups are legitimately useful compared to cloud. I’ve struggled a lot to get something useful out of the hardware I have. I realize I’m somewhat limited (16GB RX 9070), but still, it seems really far off from the kind of experience even a basic $20/month subscription gets me. Any tips anyone might have are appreciated! I’d love to be local first and would be willing to buy hardw…

A simple example. I have an older M2 Mac mini that does the OCR and visual description of all my screenshots. Screenshots are stored on my NAS. I like to screenshot things as a quick way to remember. They are things that I would not be comfortable sending a cloud provider (customer data, prototype screenshots, bank dispute details). It runs Qwen3.5:9b and glm5.2-ocr with Ollama and uses about 10GB of RAM. It automati…

> I have an older M2 Mac mini that does the OCR and visual description of all my screenshots. Screenshots are stored on my NAS.

Doesn't Apple do this already within it's OS all locally? It certainly does it for OCR and categorization.

EDIT: Also, no reason to use a generic LLM for this. This functionality exists in something like Immich (both OCR and 'context categorization'), and doesn't tie you into the Apple ecosystem either.

Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio

#164

Earlier quoted context omitted.

Banks, Biglaw, and the Pentagon all do it in the cloud. What could an individual be working on that is so secretive?

> Banks, Biglaw, and the Pentagon all do it in the cloud. In _a_ cloud: their own virtual private cloud. They also have enough power to negotiate contracts with strong privacy provisions.

This kind of stuff is available off the shelf at any major cloud provider.

Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio

#165

Earlier quoted context omitted.

[flagged]

They use the datacenters for that.

No, that's for generating the hot air used to help inflate their egos. Heating is a very separate line item on the budget.

Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio

#166

I’m curious to know if these local AI setups are legitimately useful compared to cloud. I’ve struggled a lot to get something useful out of the hardware I have. I realize I’m somewhat limited (16GB RX 9070), but still, it seems really far off from the kind of experience even a basic $20/month subscription gets me. Any tips anyone might have are appreciated! I’d love to be local first and would be willing to buy hardw…

You're limited by the manufacturer (CUDA is king, thus NVIDIA is the king right now) and your lack of VRAM will make using a useful model difficult.

I'm not surprised at all.

Context: I have a farm of DGX Sparks and several RTX 6000's, and can run very close to foundational models with ~2 sparks

Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio

#167
post #71

I’m curious to know if these local AI setups are legitimately useful compared to cloud. I’ve struggled a lot to get something useful out of the hardware I have. I realize I’m somewhat limited (16GB RX 9070), but still, it seems really far off from the kind of experience even a basic $20/month subscription gets me. Any tips anyone might have are appreciated! I’d love to be local first and would be willing to buy hardw…

I’ve been running DeepSeek 4 Flash, Qwen 27B and Qwen 9B on local hardware. They work well for coding and document review tasks. I think Qwen 9B local on a 5090 might be legitimately helpful for small task agents in omp, since it is ridiculously fast. But my motivation is that I have data that I unfortunately can’t share with 3rd parties. I have been eyeing a 512 GB Mac 5 Ultra to run full DS4 pro locally, which I ex…

I have a RTX PRO 6000 96GB when the pricing was way better than now i also have a RTX 5090 too.

What I noticed is that (1) the great local models are optimized run inference (diffusion & LLMs) well on 32GB VRAM (2) The quality of local models (esp. in diffusion) is increasing faster than the need for more VRAM - additional reason for the value of these FAST GPUs to increase!

(3) RTX PRO 6000 96GB is really great for fine tunes (ai-toolkit) :) but doesn't outperform my RTX 5090 with inference by anything significant on the good local models.

I have never run an AI job on a Mac, i also have doubts about performance and compatibilities - since the reviews almost never compare directly.

Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio

#168

Not just the high end stuff. The Neo is sold out until late September on the budget end, it seems like it is a smash for HS and college kids. I hope Apple can take all this cash and do some stability releases like they used to do, bugs around things like Family Sharing, the painful "update" to Settings App, etc could all use a lot of love.

Huh, glad I grabbed my Neo two weeks ago. It's the "top" spec version, but still a good bit less than a MBA - seemed like a pretty reasonable replacement for the M1 iPadPro it replaced (wanted to go back to a normal laptop vs tablet).

Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio

#169

Earlier quoted context omitted.

A simple example. I have an older M2 Mac mini that does the OCR and visual description of all my screenshots. Screenshots are stored on my NAS. I like to screenshot things as a quick way to remember. They are things that I would not be comfortable sending a cloud provider (customer data, prototype screenshots, bank dispute details). It runs Qwen3.5:9b and glm5.2-ocr with Ollama and uses about 10GB of RAM. It automati…

> I have an older M2 Mac mini that does the OCR and visual description of all my screenshots. Screenshots are stored on my NAS. Doesn't Apple do this already within it's OS all locally? It certainly does it for OCR and categorization. EDIT: Also, no reason to use a generic LLM for this. This functionality exists in something like Immich (both OCR and 'context categorization'), and doesn't tie you into the Apple ecosy…

[deleted]

Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio

#170

I’m curious to know if these local AI setups are legitimately useful compared to cloud. I’ve struggled a lot to get something useful out of the hardware I have. I realize I’m somewhat limited (16GB RX 9070), but still, it seems really far off from the kind of experience even a basic $20/month subscription gets me. Any tips anyone might have are appreciated! I’d love to be local first and would be willing to buy hardw…

> it seems really far off from the kind of experience even a basic $20/month subscription gets me. The $20/month subs are much stronger than the local models you can run, even with how far local models have advanced lately. The appeal of local models is that the data never leaves your network so you can feel safer putting sensitive content into it. It also feels “free” to use when you’ve already paid for the hardware…

> much stronger than the local models you can run

but depending on what you're doing, you may not need the "bleeding edge" performance

Post reply on HN