I run a lot of local models (I am always experimenting) on my 32G M2-Pro MacMini - I would love to upgrade. The financial aspects don’t work however: I can learn and experiment with what I have for local models, and I pay as I go on FireWorks.ai for open model inferencing and no matter how much I use this service my monthly bill is between $10 and $40 and much faster than any reasonable home rig. Hybrid ‘small local’…
Honest question - why are you so stuck on Macs for local inference? A 4 GPU linux box with 3090s, which are $1500 a piece right now, will blow this thing out of the water. Even 2x3090 rig will run most of the good local models like Gemma4:31b at 100+ tok/sec The VRAM of the GPUs are MUCH faster than the unified ram within Apple Silicon. The only difference is the initial model load, which takes longer from disk to VR…
A PC with similar capabilities is going to sound like a jet taking off.