Live data from Hacker News

Apple caught off guard by AI demand for Mac Mini and Mac Studio

macrumors.com

131–140 of 636 posts

Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio

#131
post #122
post #70

Earlier quoted context omitted.

You should listen to the podcast Acquired, specifically Nvidia and then Jensen Huang. They basically lucked into AI. Some researcher was using Nvidia gaming cards, and reached out to them about questions on CUDA. That email eventually turned them into a trillion dollar question.

to their credit, there was a lot of work behind "luck". Jensen showed up in person in 2017 in NEURIPS and he and likely a lot of his top brass basically sat down and read the entire conference proceedings/abstracts; there was likely a lot of work behind the scenes to behind the ML research pivot.

Yeah, The NVIDIA Way goes into a lot of detail on how and why the pivot from graphics to AI happened. This is a prime example of “you make your own luck.” Jensen engineered an organization that was primed to recognize and pounce on the next big thing, and it ended up being AI. But they saw it coming WAY in advance (like 2011/2012, not 2017) because they were explicitly on the lookout.

Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio

#132
post #68

Earlier quoted context omitted.

>I realize I’m somewhat limited (16GB RTX 9070), but still, it seems really far off from the kind of experience even a basic $20/month subscription gets me. I just ordered a new Mac Studio M5 Max 128GB $5899 ($6400 with tax) to be able to run the bigger "consumer size" models in the 70B parameter range (~96 GB). That said, I have no illusions that this expensive setup with a Qwen Flash coding LLM will be comparable t…

Serious question: why not run DGX Spark or Framework Desktop, at 30%-50% lower cost?

M5 Ultra has 4-5x the memory bandwidth of both. 1.2 TB/s memory bandwidth opens up good performance on relatively large models.

Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio

#134
post #50

Earlier quoted context omitted.

> If I had to pick a product, I'd say an affordable 32GB mac would be the sweet spot for running local models that function well like Qwen 3.8. 32GB is not enough RAM. I don't even own a device with less than 36GB at this point, and that device I only have because my employer is being cheap. 64GB is a reasonable starting point for running local LLMs + normal tasks. 128GB let's you really run most smaller models like…

You are mistaken. I'm running Qwen 3.7 28B 4bit (MLX) with a 200k context window and everything total is 32GB RSS. Is this the best? No. That's why I said the sweet spot. Getting from 16GB macs to 32GB is perhaps possible. Jumping to 64GB or 128GB as the default is simply unreasonable right now.

I assume you mean Qwen 3.8-27B? Yes, you can run this in 32GB of RAM, but it's very context limited. With KV cache compression and other techniques, it's better now than in the past, but I'd still want more RAM, personally.

EDIT to add that you need to reserve 8GB for the system if you don't want to cause problems on macOS, which means 32GB RAM = 24GB max for model + context. It takes 18-19GB to load a 4-bit quant of Qwen3.8-27B, so I'd be really surprised if you can actually get a 200k context window. You need to fit within a 24GB WSS (which is generally a more constrained RSS) to get stable performance on 32GB RAM.

Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio

#135

I’m curious to know if these local AI setups are legitimately useful compared to cloud. I’ve struggled a lot to get something useful out of the hardware I have. I realize I’m somewhat limited (16GB RX 9070), but still, it seems really far off from the kind of experience even a basic $20/month subscription gets me. Any tips anyone might have are appreciated! I’d love to be local first and would be willing to buy hardw…

A simple example. I have an older M2 Mac mini that does the OCR and visual description of all my screenshots. Screenshots are stored on my NAS. I like to screenshot things as a quick way to remember. They are things that I would not be comfortable sending a cloud provider (customer data, prototype screenshots, bank dispute details). It runs Qwen3.5:9b and glm5.2-ocr with Ollama and uses about 10GB of RAM. It automati…

What is your M2's total memory? I find this application really interesting.

Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio

#136

Earlier quoted context omitted.

> AI subscriptions are too highly subsidized right now I've been running into annoying limits with Claude recently. It gives me like 5 questions over the course of 15 mins and then tells me to wait 5 hours. When companies can change things up to make the base subscription nearly useless (the last question always gets messed up, too), then you realize the value of owning your own infrastructure.

On the $200/mo plan I have never hit a five hour limit, and I struggle to use my full credits each week. $200/month is vastly cheaper than owning and operating comparable hardware.

You're right that $200/mo is much cheaper than comparable infrastructure. OTOH, you don't get to have a computer that can also be used for other applications, or which works when the internet is down. Also, you remain tethered to whatever pricing the AI companies want to charge. If AI pricing goes like Uber/Lyft did when the VC cash ran out, then we'll be paying much more in a few years' time. We could look back and think "I wish I'd bought my own setup back in 2026" if it's going to be inevitable.

And all this before you get into privacy/security/compliance stuff.

Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio

#137
post #50

Earlier quoted context omitted.

> If I had to pick a product, I'd say an affordable 32GB mac would be the sweet spot for running local models that function well like Qwen 3.8. 32GB is not enough RAM. I don't even own a device with less than 36GB at this point, and that device I only have because my employer is being cheap. 64GB is a reasonable starting point for running local LLMs + normal tasks. 128GB let's you really run most smaller models like…

You are mistaken. I'm running Qwen 3.7 28B 4bit (MLX) with a 200k context window and everything total is 32GB RSS. Is this the best? No. That's why I said the sweet spot. Getting from 16GB macs to 32GB is perhaps possible. Jumping to 64GB or 128GB as the default is simply unreasonable right now.

Memory used : 38GB, and I haven't even started a LLM nor podman, I always fight with memory when using LLM on my mac with 48gb.

And I don't remember to have been able to have pushed to 200k context Qwen 3.6. 3.8 is running on my RTX 5090.

Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio

#138

Earlier quoted context omitted.

It also takes some load off the AI data centers. IDK if that might be a concern for Apple or their AI partners.

It worsens the supply crunch, no? A unit you use sparingly vs that memory going into a GPU that serves many more people.

Those will use different wafers, so unless that memory is allocated for unified memory vs gpu HBM it won't make a difference.

Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio

#139

Earlier quoted context omitted.

> You have good enough hardware to run good models comparable with Gemini and ChatGPT. That is at best misleading and at worst outright misinformation.

If you have 2TB of VRAM you can’t run one of the big models which are comparable?

The post was replying to someone with 16GB. (And also: no, even the best open weight models are not as good as what you can use on your ChatGPT subscription. They’ve gotten a lot better, but not that much better.)

Re: Apple caught off guard by AI demand for Mac Mini and Mac Studio

#140
post #68

I’m curious to know if these local AI setups are legitimately useful compared to cloud. I’ve struggled a lot to get something useful out of the hardware I have. I realize I’m somewhat limited (16GB RX 9070), but still, it seems really far off from the kind of experience even a basic $20/month subscription gets me. Any tips anyone might have are appreciated! I’d love to be local first and would be willing to buy hardw…

>I realize I’m somewhat limited (16GB RTX 9070), but still, it seems really far off from the kind of experience even a basic $20/month subscription gets me. I just ordered a new Mac Studio M5 Max 128GB $5899 ($6400 with tax) to be able to run the bigger "consumer size" models in the 70B parameter range (~96 GB). That said, I have no illusions that this expensive setup with a Qwen Flash coding LLM will be comparable t…

> Why did I initially spend the extra $1600 if I knew ahead of time that it wasn't as good as cloud AI? Because I thought I could use some local LLM for the easy tasks or when I hit cloud rate limits.

The maths don't check. With Deepseek Flash one goes a very long way with 1600$ - even 10$/month, for easy jobs, are more than 13 years, and at a higher quality.

Post reply on HN