Sure. Wealth has always given early adopters and first movers wings.
As hardware specs get better, and model providers keep making different classes of mistakes that alienate users, more people will opt for or support local models. I tend to take Schneiers' "Attacks only improve" philosophy and apply it to damn near everything like this. Inference will get more expensive, but IMO cloud based inference is at or near the maximum that users can tolerate, across multiple spectra. There may be a marginal increase in cost, but as the different metrics between what is affordable for rapidly scalable, cloud based inference and the lag of local inference with open models and expensive hardware converge, prices will come down.
I am not the right person to say what the peak is going to look at, but when the companies with (practically speaking) unlimited compute resources and money are starting to take a good hard look, there is going to be a drive towards efficiency. I hear grumbling at work, and the appreciation from my leaders when I show up with a tool that shows how I reduced my token costs, or I keep asking the questions of what the token cost (real, and internal billing) are for tools my peers write are.
I also have absolutely eye-watering personal AI bills that are starting to make buying a better tier of local inference gear look more affordable, even with inflated hardware prices. The current generation of open models are very effective, and paying for multiple $200 dollar a month subs is not feasible longer term, and as one example, in one month, if I was paying enterprise token rates, I would have burned through $8000+ token budgets on one of those accounts, doing actual, practical useful work, it makes more sense to buy some much better hardware than to pay for one-time inference costs.