Live data from Hacker News

Ask HN: MacBook vs. Dedicated GPU for LLM

news.ycombinator.com

51–60 of 76 posts

Re: Ask HN: MacBook vs. Dedicated GPU for LLM

#51
post #2

MacBooks with their unified memory behave like a slow GPU with enormous amount of video RAM. So you can run large smart models slowly. Dedicated GPUs have less video RAM so can run smaller less smart models quickly.

> MacBooks with their unified memory behave like a slow GPU with enormous amount of video RAM. So you can run large smart models slowly. With the model using MLX the speed increase is night and day. Even non-MLX is good. You also don't have the transfer costs related to moving CPU data into the GPU.

[dead]

Re: Ask HN: MacBook vs. Dedicated GPU for LLM

#52
post #15

Earlier quoted context omitted.

Macbook M5 64GB - can run gemma-4-26b-a4b-it-4bit and Qwen3.6-35B-A3B-4bit at about 1500 tps prefix and 45 tps decode on contexts up to 100K tokens using MLX. It's faster than Claude. I was really surprised, chat quality is also similar to Claude for gemma4. Agentic works but does not compare to cloud models, you can still make agents where top level is code.

sorry but asking again: how much memory is actually useable by gpu in macbook? as it is shared(os and apps also have to use same memory)? and it is different than dedicated gpu memory?

Rule of thumb is about 70-80%.

Re: Ask HN: MacBook vs. Dedicated GPU for LLM

#53
post #7

My opinion is that you should wait for 6-12 months before making a purchase either way. Open weight models are getting good. With GLM 5.2 now chasing Opus, I'm very excited to see a smaller model's distillation. Plus, the OLED MacBook Pro should be released by then.

The OLED touchscreen MacBook is rumoured to be called MacBook Ultra now, and it has been delayed quite a few times. I will probably cost the same than a decent bike.

MacBook Ultra Expensive (when equipped with 256 or 512 Gb of ram)

Re: Ask HN: MacBook vs. Dedicated GPU for LLM

#55
post #7

My opinion is that you should wait for 6-12 months before making a purchase either way. Open weight models are getting good. With GLM 5.2 now chasing Opus, I'm very excited to see a smaller model's distillation. Plus, the OLED MacBook Pro should be released by then.

The OLED touchscreen MacBook is rumoured to be called MacBook Ultra now, and it has been delayed quite a few times. I will probably cost the same than a decent bike.

> a decent bike

I bought a decent bike a few months ago for a few hundred bucks. I used it for commuting and for when I went to the park.

Of course, it depends on what kind of bike you meant and what you consider decent. Hopefully that illustrates the quality of that comparison.

Re: Ask HN: MacBook vs. Dedicated GPU for LLM

#56
So a lot depends on your specific use case but mid-sized open weight models are pretty actually good now, so this is realistic [1]

The first question to ask is does your use case require handling personal or sensitive data.

If you're using the LLM for OpenClaw or you want to handle sensitive or medical data, a local model generally is necessary.

if it's not so sensitive - Cloud providers with some sort of user agreement guarantee on not using your data for training would be the next bet. I personally generally use Gemini or Sonnet as my cloud backup. As I understand, OpenAI, Cloudflare (which bought replicate) and Qwen also seem to provide such guarantees and make SOTA models available. Others like DeepSeek seem to have an opt-out setting. Open router & co I avoid except for benchmarking models with public or dummy data as there is absolutely zero guarantee or ability to enforce terms on providers where your data might be sent.

Gemini and Anthropic (and OpenAI) tend to be expensive - it's very easy to run up 15 dollars a day or so bills which puts you solidly in 1 year pay out on Mac Mini territory - at this point I decided to buy. Gemini Flash Lite 3.1 is however surprisingly good value.

the next question is Mac or CUDA. If your expected use is serving LLM models for inferences, the latest large memory Macs give pretty good inference speed (better than DGX Spark) at a reasonable cost - I think there offer much better value than CUDA if the only use case is LLM inference & harnesses.

if you plan to also fine tune models, experiment with other types of ML on GPUs, do computer vision stuff etc. the development tooling on CUDA is far in advance of all other platforms.

Lastly if you choose CUDA, the question is GB10 family (DGX Spark - cluster able with 128Gb RAM et all) or dedicated GPUs workstations. What I found is practically any serious models weighs in requiring at least 96GB VRAM - Antirez's 2 bit quant of Deepseek 4 flash (my current daily driver) [2] , the Qwen 3.5 122B A10B 4-bit quant, the Qwen 3.6 27B Dense and 35B A3B 8 but quants etc. So you're well out of the consumer GPU territory into 1 or more RTX 6000 Pros or Data center grade devices. Yes you can try to hack away with multiple consumer cards or SSD streaming but it's very fiddly and you probably have better things to do with your life.

The GB10 system - which I ultimately went with - is certainly much cheaper and can be clustered through the Special NVLink cable to get 256, 384 or 512 GB setups but comes with severely constrained bandwidth. The Pro GPUs blast these out of water on performance but are expensive.

Lastly, renting a cloud GPU machine doesn't really make sense except to run already debugged fine tuning workloads. You'll probably spend at least 4 dollar an hour for sufficient capacity which if it's personal use, will mostly sit idle.

1. https://srinathh.medium.com/mid-size-local-models-are-now-co...

2. https://github.com/antirez/ds4

Re: Ask HN: MacBook vs. Dedicated GPU for LLM

#57

Earlier quoted context omitted.

The OLED touchscreen MacBook is rumoured to be called MacBook Ultra now, and it has been delayed quite a few times. I will probably cost the same than a decent bike.

> a decent bike I bought a decent bike a few months ago for a few hundred bucks. I used it for commuting and for when I went to the park. Of course, it depends on what kind of bike you meant and what you consider decent. Hopefully that illustrates the quality of that comparison.

My commuting bike is okay and in this price range too.

The MacBook Neo is also a good enough laptop.

In my opinion, a decent bike is something that wouldn’t limit you in races. No need to spend enormous amounts of money for marginal gains, but something that would do the job well. That’s an order of magnitude more expensive.

Re: Ask HN: MacBook vs. Dedicated GPU for LLM

#58

Earlier quoted context omitted.

> a decent bike I bought a decent bike a few months ago for a few hundred bucks. I used it for commuting and for when I went to the park. Of course, it depends on what kind of bike you meant and what you consider decent. Hopefully that illustrates the quality of that comparison.

My commuting bike is okay and in this price range too. The MacBook Neo is also a good enough laptop. In my opinion, a decent bike is something that wouldn’t limit you in races. No need to spend enormous amounts of money for marginal gains, but something that would do the job well. That’s an order of magnitude more expensive.

So the MacBook Neo costs the same as a decent bike? What makes you think the Ultra will be in that range as well?

Re: Ask HN: MacBook vs. Dedicated GPU for LLM

#59

Earlier quoted context omitted.

My commuting bike is okay and in this price range too. The MacBook Neo is also a good enough laptop. In my opinion, a decent bike is something that wouldn’t limit you in races. No need to spend enormous amounts of money for marginal gains, but something that would do the job well. That’s an order of magnitude more expensive.

So the MacBook Neo costs the same as a decent bike? What makes you think the Ultra will be in that range as well?

I wouldn’t race on my commuter bike or work professionally on a MacBook Neo.

I guess it could start around 3000€

Re: Ask HN: MacBook vs. Dedicated GPU for LLM

#60

Earlier quoted context omitted.

but still it can run handsome models

Latest MBP goes up to 128GB of memory.

The M5 Max's GPU is slower than a single RTX 3090, though.

Both machines will be stuck in the 30-50GB model range, but the 3090 would have faster token prefill and faster decode speeds (614GB/s on M5 Max vs 936GB/s on 1x 3090).

Post reply on HN