Well, I don't see a value issue of using Qwen3.6 27B vs Sonnet 4.6 (not sure about 5 yet)
I still have to use GHCP at work, and I self-host at home, and aside from the fact self-hosting also forces you to tinker, optimize, etc. - there's not a huge difference in my end result in end user results. I spent quite a bit of time trying to optimize llamacpp and compare 35b to 27b, etc. I don't compare models that much at work.
I guess the other part of it is I didn't really know much about cheaper cloud models, but I was attracted to the idea of no longer renting against Claude code, etc. I figured if I could run something functionaly similar from my bedroom on a normal outlet, then all this talk about data centers needing to be built everywhere in the news cycle is obviously just plain stupidity and hype.
It appears I'm using about 20.4/7.6 million in/out tokens a month, or on open router, about $20/month.
That puts $1350 at a 5-6 year break even (thanks to cheap electricity), I guess. Beyond that, running on localhost as a nice feature of 0 no latency when doing rapid tool calling