Not going to bother with this until the temporal cohesion issue is solved. The results are cool, but the variety in frames makes it look like a very specific and distracting art style instead of true animation.
Did you scroll down and see the different methods used, specifically method 5? Not perfect, but getting pretty close.
I suspect depth maps (SD2 already supports them) could be used to achieve that in the future.
I wonder if a diffusion model could accept an "onion skin" noise, so the transitions between frames would be less jarring. Can someone with more knowledge than me explain what's the most promising approach here?