If decode is becoming so complicated and expensive the hardware can't handle it, why not just go full neural, send latents, and run decode on tensor cores? The answer is probably the same as for why not AV2 everything; a lot of hardware couldn't support it today. But in 10 years? It seems we're running up against fundamental limits of human-engineered video codecs at this point. There might be a lesson in there.
And it's not really hardware hitting limits, it's specifically software decoding on somewhat weaker machines.