Earlier quoted context omitted.
Inefficient per watt, yes, but local inference capacity is greatly underutilized in aggregate. If a model can run on a machine that already exists, that's a bunch of additional chips that don't need to be built.
You picked up on what I didn't write existing machines don't have to be built, you just use them. It occurred to me much later that local machines could become a shared resource of a small co-op. Or it only runs off solar or energy harvested at times of excess generation (too much daytime solar, lots of wind, or overnight when demand naturally drops). Exploring how to partition inference across many machines and shed…
It's also resilience. I can still do AI stuff if the internet goes down.