I like that this one shows some "fails", and not just the top of the top results: For example, the surfer is surfing in the air at the end: https://cdn.openai.com/tmp/s/prompting_7.mp4 Or this "breaking" glass that does not break, but spills liquid in some weird way: https://cdn.openai.com/tmp/s/discussion_0.mp4 Or the way this person walks: https://cdn.openai.com/tmp/s/a-woman-wearing-a-green-dress-a... Or wherever…
Video generation models as world simulators
31–40 of 171 posts
Re: Video generation models as world simulators
#321. High quality video or image from text 2. Taking in any content as input and generating forwards/backwards in time 3. Style transformation 4. Digital World simulation!
Re: Video generation models as world simulators
#33Re: Video generation models as world simulators
#34Reasoning, logic, formal systems, and physics exist in a seemingly completely different, mathematical space than pure video.
This is just a contrived, interesting viewpoint of the technology, right?
Re: Video generation models as world simulators
#35I think people might be missing what this enables. It can make plausible continuations of video, with realistic physics. What happens if this gets fast enough to work _in real time_. Connect this to a robot that has a real time camera feed. Have it constantly generate potential future continuations of the feed that it's getting -- maybe more than one. You have an autonomous robot building a real time model of the wor…
Imagine where you want to be (eg, “I scored a goal!”) from where you are now, visualize how you’ll get there (eg, a trick and then a shot), then do that.
Re: Video generation models as world simulators
#36AlphaGo and AlphaZero were able to achieve superhuman performance due to the availability of perfect simulators for the game of Go. There is no such simulator for the real world we live in (although pure LLMs sort of learn a rough, abstract representation of the world as perceived by humans.) Sora is an attempt to build such a simulator using deep learning. “Our results suggest that scaling video generation models is…
Re: Video generation models as world simulators
#37I can't wait to play with this but I can't even imagine how expensive it must be. They're training in full resolution and can generate up to a minute of video.
Seeing how bad video generation was, I expected it would take a few more years to get to this but it seems like this is another case of "Add data & compute"(TM) where transformers prove once again they'll learn everything and be great at it
Re: Video generation models as world simulators
#38I like that this one shows some "fails", and not just the top of the top results: For example, the surfer is surfing in the air at the end: https://cdn.openai.com/tmp/s/prompting_7.mp4 Or this "breaking" glass that does not break, but spills liquid in some weird way: https://cdn.openai.com/tmp/s/discussion_0.mp4 Or the way this person walks: https://cdn.openai.com/tmp/s/a-woman-wearing-a-green-dress-a... Or wherever…
The hyper realistic and plausible movement of the glass breaking makes this bizarrely fascinating. And it doesn’t give me the feeling of disgust the motion in the more primitive AI models did
Re: Video generation models as world simulators
#39Earlier quoted context omitted.
Our ability to build somewhat convincing simulations of thing has never been a proof of living in a simulation…
i mean everyone's mind builds a convincing internal simulation of reality and it's so good that most people think they're directly experiencing reality.
Re: Video generation models as world simulators
#40So this is why they haven't shown Will Smith eating spaghetti.
> These capabilities suggest that continued scaling of video models is a promising path towards the development of highly-capable simulators of the physical and digital world
This is exciting for robotics. But an even closer application would be filling holes in gaussian splatting scenes. If you want to make a 3D walkthrough of a space you need to take hundreds to thousands of photos with seamless coverage of every possible angle, and you're still guaranteed to miss some. Seems like a model this capable could easily produce plausible reconstructions of hidden corners or close up detail or other things that would just be holes or blurry parts in a standard reconstruction. You might only need five or ten regular photos of a place to get a completely seamless and realistic 3D scene that you could explore from any angle. You could also do things like subtract people or other unwanted objects from the scene. Such an extrapolated reconstruction might not be completely faithful to reality in every detail, but I think this could enable lots of applications regardless.