Earlier quoted context omitted.
If models become more efficient we will move more of the work to local devices instead of using SaaS models. We’re still in the mainframe era of LLM.
The hyperscalers do not want us running models at the edge and they will spend infinite amounts of circular fake money to ensure hardware remains prohibitively expensive forever.
If that's the plan (there is no plan) then it expires at some point, because it's a spiral and such spirals always bottom out.