This seems to be a lot of effort to rationalize the surprisingly small performance increase from M2 to M3. Initially the assumption was that M2 to M3 would be a bigger step than M1 to M2, not a smaller one. Perhaps TSMC 3nm is showing the limits of scaling?
I don't think we're going to see the performance increases past the ~10-20% each fab cycle now.
Performance per watt and having dedicated HW for video decoding / other tasks(AI?) I think might end up being the answer to the continuation of Moore's law. Especially considering that we are in the era of heterogenous computing.