32B is a good choice of size, as it allows running on a 24GB consumer card at ~4 bpw (RTX 3090/4090) while using most of the VRAM. Unlike llama 3.1, which had 8b, 70B (much too big to fit), and 405B.
what do you mean? I can easily run 70b on my macbook. Fits easily.
Does your MacBook really have a 24GB VRAM consumer (GPU) card?