Earlier quoted context omitted.
Serving barely useful GLM 5.2 costs what? $15k? Actually useful is like $50k? You’ll never recoup the cost unless you ‘locally’ means ‘inference provider is not the model provider’?
The high costs are necessary for high speed. When a low speed of the order of one token per second is accepted, any open weights LLM can be run on an ordinary PC (with the weights read from SSDs) and the cost becomes negligible. Such a low speed would be annoying for a chat, but I do not believe that it is "barely useful" for a coding assistant. There are plenty of tasks for which it is fine to get results some hours…
So maybe for a hobby project this is fine, but for something you have to take to market and compete with... I think it'd be a really rough sell.
EDIT: also, just to be clear: if there was a practical path to using local AI, I'd take it in a heartbeat. I hope it gets to the point that it's better to use local than paying someone $200/mo. But right now, that $200/mo is the clear best option. I get making compromises for ideology but the compromises are too big for me right now.