I think this sort of proves out that video is a fair bit further away than just text-based reasoning becoming mass market. When you use just text the mind is able to interpolate away and/or fill in some of the uncanny parts. With video, it's just not possible -- the uncanny valley is novel at first, but boring in the long run. Video will have a much much higher bar to pass than text and I think it will take some time…
I am not so sure about that. The real hurdle at this point seems to be getting something like Dall-E to produce multiple angles and poses of a scene while maintaining characteristics of the elements in the scene. At that point it is an interpolation problem. I admit I am not an expert on these, but these do not seem insurmountable. I would not be surprised if we see something that can make very convincing short videos in less than a year. At that point, since a "short video" is really a scene, we are basically at film stage.