This logic seems mad. If people only need SLMs then hyperscalers can also centrally host higher-efficiency models, and still gain efficiencies of scale and convenience over hosting locally.
The efficiency doesn't matter as long as it's good enough, because you already have the computer.