Transformer models are quickly becoming a commodity, and I suspect in time we'll all be running them locally. Even now, you can run something pretty useful on a 16 GB graphics card, and I suspect a decade into the future, entry level hardware will be running better models than high-end graphics cards can run now, as entry-level hardware gets better and models get more efficient. It doesn't mean hosted frontier models…
I doubt it. The play seems to be: lock what was once commodity compute up into datacenters depriving us regular folk of it, then sell it back to us on subscription. Even if my #NeverSubscribe movement succeeds, all that misdirected hardware [into datacenters] won't likely be practical for home use.