Earlier quoted context omitted.
Honest question - why are you so stuck on Macs for local inference? A 4 GPU linux box with 3090s, which are $1500 a piece right now, will blow this thing out of the water. Even 2x3090 rig will run most of the good local models like Gemma4:31b at 100+ tok/sec The VRAM of the GPUs are MUCH faster than the unified ram within Apple Silicon. The only difference is the initial model load, which takes longer from disk to VR…
A Mac mini running an LLM is quiet A PC with similar capabilities is going to sound like a jet taking off.
Not at all. Airflow with big fans is quiet. What I do hear is coil-whine. In fact my PC is quieter than my Macbook when both are running top speed. But one has 4000 AI TOPS.