Video generation models as world simulators
1–10 of 171 posts
Re: Video generation models as world simulators
#2Re: Video generation models as world simulators
#3I find it wild that this model does not have explicit 3D prior, yet learns to generate videos with such 3D consistency, you can directly train a 3D representation (NeRF-like) from those videos: https://twitter.com/BenMildenhall/status/1758224827788468722
Re: Video generation models as world simulators
#4I find it wild that this model does not have explicit 3D prior, yet learns to generate videos with such 3D consistency, you can directly train a 3D representation (NeRF-like) from those videos: https://twitter.com/BenMildenhall/status/1758224827788468722
Re: Video generation models as world simulators
#5The robot examples are very underwhelming, but the people and background people are all very well done, and at a level much better than most static image diffusion models produce. Generating the same people as the interact with objects is also not something I expected a model like this to do well so soon.
Re: Video generation models as world simulators
#6Edit, changed the links to the direct ones!
Re: Video generation models as world simulators
#7 “Our results suggest that scaling video generation models is a promising path towards building general purpose simulators of the physical world.”
General, superhuman robotic capabilities on the software side can be achieved once such a simulator is good enough. (Whether that can be achieved with this approach is still not certain.)Why superhuman? Larger context length than our working memory is an obvious one, but there will likely be other advantages such as using alternative sensory modalities and more granular simulation of details unfamiliar to most humans.
Re: Video generation models as world simulators
#8Re: Video generation models as world simulators
#9Damn, even minecraft videos being simulated, this is crazy to see from OpenAI. Edit, changed the links to the direct ones! https://cdn.openai.com/tmp/s/simulation_6.mp4 https://cdn.openai.com/tmp/s/simulation_7.mp4
example video links from TFA:
Re: Video generation models as world simulators
#10I find it wild that this model does not have explicit 3D prior, yet learns to generate videos with such 3D consistency, you can directly train a 3D representation (NeRF-like) from those videos: https://twitter.com/BenMildenhall/status/1758224827788468722
The crazy thing is that they do it by prompting the model to in paint a chrome sphere into the middle of the image to reflect what is behind the camera! The model can interpret the context and dream up what is plausibly in the whole environment.