Everything I've seen makes me suspect that models have continually got more efficient to serve.
The strongest evidence is that the models I can run on my own laptop got massively better over the last three years, despite me keeping the same M2 64GB machine without upgrading it.
Compare original LLaMA from 2023 to gpt-oss-20b from this year - same hardware, huge difference.
The next clue is the continuing drop in API prices - at least prior to the reasoning rush of the last few months.
One more clue: o3. OpenAI's o3 had a 80% price drop a few months ago which I believe was due to them finding further efficiencies in serving that model at the same quality.
My hunch is that there are still efficiencies to be wrung out here. I think we'll be able to tell if that's not holding if API prices stop falling over time.