This has been coming for a long time and it is why I use local models only. I'm willing to give up capabilities in exchange for being able to trust that whatever biases may exist in the models I do use remain static and predictable.
How do you do it? Do you host on your hardware or do you use cloud-based providers for open models?
Funny enough the mac has almost the same processor as my iPhone 16 Pro, so its just a RAM constraint, and of course PrivateLLM does not let you host an API.
An M4 Pro would do much better do to the increase in RAM and GPU size.