My opinion is that you should wait for 6-12 months before making a purchase either way. Open weight models are getting good. With GLM 5.2 now chasing Opus, I'm very excited to see a smaller model's distillation. Plus, the OLED MacBook Pro should be released by then.
This is my opinion too. Even if you buy hardware like a cluster of 8xGB10s or 4 A100s, they'll still be slow and a little dumber than what you're used to. We need to wait a little for better hardware. Lots of companies are pushing the frontier, so hopefully it'll come very soon.
Competition and innovation will hopefully make the bubble pop, and we'll get reasonably priced local hardware to run very intelligent models. Something like Talaas with GLM 5.2 would be pretty cool. Or Apple printing the latest model onto hardware—it would give a new reason to buy a new Mac every year (a new ai model with every new version).
MacBooks with their unified memory behave like a slow GPU with enormous amount of video RAM. So you can run large smart models slowly. Dedicated GPUs have less video RAM so can run smaller less smart models quickly.
Do Mac Pros provide more headroom? noob here, noob questions
My opinion is that you should wait for 6-12 months before making a purchase either way. Open weight models are getting good. With GLM 5.2 now chasing Opus, I'm very excited to see a smaller model's distillation. Plus, the OLED MacBook Pro should be released by then.
This is my opinion too. Even if you buy hardware like a cluster of 8xGB10s or 4 A100s, they'll still be slow and a little dumber than what you're used to. We need to wait a little for better hardware. Lots of companies are pushing the frontier, so hopefully it'll come very soon. Competition and innovation will hopefully make the bubble pop, and we'll get reasonably priced local hardware to run very intelligent models…
The hardware is here today for people prepared to tolerate mild amounts of latency. It’s easy to forget that computing tasks used to often take major amounts of time - rendering an audio file, rendering a video, transcoding – all kinds of tasks took minutes or even hours of the computer spinning its fans on maximum just to deliver the result. AI and agentic AI and diffusion is the next round of that - trading a small bit of your waiting time for phenomenal power. The datacentre builders trying to get you hooked on instant responses on the LLM platforms have made you think that a “good” AI responds instantly and completely interactively - they can still be brilliant with a bit of delay. And having a competent agent doing things on my local machine, it doesn’t really matter if it takes ten minutes or an hour or six hours to complete a task while I’m out doing other things.
If you want a massive MacBook anyway then it's great. They are decent for local LLMs, awesome for local image models and it's a MacBook so AppleCare+ has your back. IMO it's a no brainer if you wanted a MacBook anyway but it's a poor choice if your reason to buy it is to run LLMs.
I agree. To run an acceptable model (e.g. Qwen/Qwen3.6-27B or google/gemma-4-31B) with a good quantization (minimum Q5) with a good context size (min 64k) you could buy 2 or even 3 GTX 5060 16GiB VRAM for ~550$ each. Fyi the much faster MoE models were useless for my usecases - e.g not able to correctly identify me/I/you, endless thinking loops, etc.
I'm currently running those models using an RTX 5070 12GiB + RTX 5060 16GiB + RTX 3060 12GiB with a 96k context size with MTP/speculative decoding and I'm quite happy (the 5070 is about 4x faster than the 3060, the 5060 is inbetween them so about 2x faster than a 3060).
It’s kind of amazing how steadily this question is asked in every forum where it can be asked. Kind of amazing that the answers previously given can’t reach the next person who’s going to ask it.
Macbook M5 64GB - can run gemma-4-26b-a4b-it-4bit and Qwen3.6-35B-A3B-4bit at about 1500 tps prefix and 45 tps decode on contexts up to 100K tokens using MLX. It's faster than Claude. I was really surprised, chat quality is also similar to Claude for gemma4. Agentic works but does not compare to cloud models, you can still make agents where top level is code.
sorry but asking again: how much memory is actually useable by gpu in macbook? as it is shared(os and apps also have to use same memory)? and it is different than dedicated gpu memory?
You can adjust the percentage available both on the MacOS side and how much the model uses.
Macbook M5 64GB - can run gemma-4-26b-a4b-it-4bit and Qwen3.6-35B-A3B-4bit at about 1500 tps prefix and 45 tps decode on contexts up to 100K tokens using MLX. It's faster than Claude. I was really surprised, chat quality is also similar to Claude for gemma4. Agentic works but does not compare to cloud models, you can still make agents where top level is code.
sorry but asking again: how much memory is actually useable by gpu in macbook? as it is shared(os and apps also have to use same memory)? and it is different than dedicated gpu memory?
This is my opinion too. Even if you buy hardware like a cluster of 8xGB10s or 4 A100s, they'll still be slow and a little dumber than what you're used to. We need to wait a little for better hardware. Lots of companies are pushing the frontier, so hopefully it'll come very soon. Competition and innovation will hopefully make the bubble pop, and we'll get reasonably priced local hardware to run very intelligent models…
The hardware is here today for people prepared to tolerate mild amounts of latency. It’s easy to forget that computing tasks used to often take major amounts of time - rendering an audio file, rendering a video, transcoding – all kinds of tasks took minutes or even hours of the computer spinning its fans on maximum just to deliver the result. AI and agentic AI and diffusion is the next round of that - trading a small…
Hmm, I have access to A100s and a GB10, but if I use the models hosted there to code, I waste a lot of time waiting for answers and correcting errors. The amount of work I get done thanks to the quality and speed of frontier hosted models let me be insanely productive and have a lot of free time. I could use the slow local setup, but at what price?
Around February or march I started looking into hardware options to help me start learning about training models and working with them. My budget was limited and an apple refurbished 32 gb Mac mini was far and away the best option for my budget. I wish it was faster but I can let it run 24/7 with no noise and minimal power draw. I just arrange long running tasks for when am asleep or at work. Then as a huge plus I have an awesome daily driver machine for whatever else I want to do
My opinion is that you should wait for 6-12 months before making a purchase either way. Open weight models are getting good. With GLM 5.2 now chasing Opus, I'm very excited to see a smaller model's distillation. Plus, the OLED MacBook Pro should be released by then.
This is my opinion too. Even if you buy hardware like a cluster of 8xGB10s or 4 A100s, they'll still be slow and a little dumber than what you're used to. We need to wait a little for better hardware. Lots of companies are pushing the frontier, so hopefully it'll come very soon. Competition and innovation will hopefully make the bubble pop, and we'll get reasonably priced local hardware to run very intelligent models…
The racks we're deploying are effectively GB300 NVL72s: 72 Blackwell Ultra GPUs 36 Grace CPUs, 20.7TB of unified HBM3e.
Works out to about 1.1exaflops of fp4. Networking is 800gbps.