Earlier quoted context omitted.
32GB of fast unified memory is enough for Qwen 3.8 27B. - 16GB for the weights at Q4 - 9GB for the full 256K context at Q8 - 7GB spare for overhead and system. The problem is that these Macs have 32GB of slow unified memory. Edit: I'm thinking of a headless Mac mini, if you meant running it on the same machine you're using of course you'll need more memory, but LLMs are best served from a headless server so that's wh…
Is this for setup for agentic coding? Why not also run the IDE compiler etc... on the same machine to use those CPU cores as well?
32gb of unified memory is enough enough for system to be used for anything other than LLM generation.