Earlier quoted context omitted.
The big question for local LLMs is whether there is a 100 tok/s model which requires less than 16 GB of memory and is competitive on most tasks with the cloud models. There is some signal that this is possible through both hardware innovation and training/data improvements. Cloud models have their own constraints - I can’t have opus4.8 spend 4 hours on a deep research question I had in the shower without spending mon…
if you do the electricity math you'll see that you pay more on local models while getting less (local is more heavily quantized) compared with OpenRouter. I'm not talking local Gemma/Qwen vs cloud Opus, but against OpenRouter same Gemma/Qwen there are reasons to run local - privacy, availability, but cost is not one of them
Now, this all brings the upfront costs way up, the solar panels are cheap, its all the rest around them that tends to cost money.