"Spatial audio that supports dynamically computed surround sound for arbitrarily many speakers and headphones."
Still covers only a small arbitrary constant factor of data that is already very small by modern standards.
Note that surround sound data is, IIRC, already not "twice" the size of stereo.
And your video points 1 and 2 are also still at most only small constant factors of increase over what we already have, with 3 potentially being a compression technique.
Video is nearing its apex; sound is pretty much already there.
There's actually a maximum rate at which our senses can convey information to our brains; any use of data beyond that rate is literally impossible and anything carried beyond that is wasted. Even a full sensorium just isn't that much larger than what we already have. We are, after all, talking about technologies that are in the same ballpark as the maximum theoretical data density that human brains can have, and in practice the memristor storage is going to be much higher. It should not be surprising that it's very difficult to truly "use" all that storage.