Live data from Hacker News

Video generation models as world simulators

openai.com

71–80 of 171 posts

Re: Video generation models as world simulators

#71

I think people might be missing what this enables. It can make plausible continuations of video, with realistic physics. What happens if this gets fast enough to work _in real time_. Connect this to a robot that has a real time camera feed. Have it constantly generate potential future continuations of the feed that it's getting -- maybe more than one. You have an autonomous robot building a real time model of the wor…

What I find interesting is that b/c we have so much video data, we have this thing that can project the future in 2d pixel space. Projecting into the future in 3d world space is actually what the endgame for robotics is and I imagine depending on how complex that 3d world model is, a working model for projecting into 3d space could be waaaaaay smaller. It's just that the equivalent data is not as easily available on…

That's what estimation and simulation is for. Obviously that's not what's happening in TFA but it's perfectly plausible today.

Not sure how people are concluding that realistic physics is feasible operating solely in pixel space, because obviously it can't and anyone with any experience training such models would recognize instantly the local optimum these demos represent. The point of inductive bias is to make the loss function as convex as possible by inducing a parametrization that is "natural" to the system being modeled. Physics is exactly the attempt to formalize such a model borne of human cognitive faculties and it's hard to imagine that you can do better with less fidelity by just throwing more parameters and data at the problem, especially when the parametrization is so incongruent to the inherent dynamics at play.

Re: Video generation models as world simulators

#72

Earlier quoted context omitted.

Except there is always an original at the root. There’s no way to prove that’s not us.

The root world can spawn many simulations and simulations can be spawned within simulations. It becomes far more likely that we exist in a simulated world than in the root world.

The only thing about the branching simulations is they are likely simplified approximations. There’s no reason it doesn’t nest and that the approximations can observe their approximations of the prior level is strictly more complex than can be observed in the simulation. That should be fundamentally impossible meaning any branch can’t know if they’re the root or the branch, only that they create a branch.

Re: Video generation models as world simulators

#73

Earlier quoted context omitted.

Our ability to build somewhat convincing simulations of thing has never been a proof of living in a simulation…

i mean everyone's mind builds a convincing internal simulation of reality and it's so good that most people think they're directly experiencing reality.

Buddhist insight meditation actually proves that’s not true, fwiw.

Re: Video generation models as world simulators

#74

I think people might be missing what this enables. It can make plausible continuations of video, with realistic physics. What happens if this gets fast enough to work _in real time_. Connect this to a robot that has a real time camera feed. Have it constantly generate potential future continuations of the feed that it's getting -- maybe more than one. You have an autonomous robot building a real time model of the wor…

You're talking about an agent with a world model used for planning. Actually generating realistic images is not really needed as the world model operates in its own compressed abstraction. Check out V-Jepa for such a system: https://ai.meta.com/blog/v-jepa-yann-lecun-ai-model-video-jo...

V-Jepa is actually super impressive. I have nothing but respect for Yann LeCun & his team, they really have been on a rampage lately.

Re: Video generation models as world simulators

#75

I think people might be missing what this enables. It can make plausible continuations of video, with realistic physics. What happens if this gets fast enough to work _in real time_. Connect this to a robot that has a real time camera feed. Have it constantly generate potential future continuations of the feed that it's getting -- maybe more than one. You have an autonomous robot building a real time model of the wor…

> Connect this to a robot that has a real time camera feed. Have it constantly generate potential future continuations of the feed that it's getting

There was that article a few months ago about how basically that's what the cerebellum does.

Re: Video generation models as world simulators

#76

I think people might be missing what this enables. It can make plausible continuations of video, with realistic physics. What happens if this gets fast enough to work _in real time_. Connect this to a robot that has a real time camera feed. Have it constantly generate potential future continuations of the feed that it's getting -- maybe more than one. You have an autonomous robot building a real time model of the wor…

What I find interesting is that b/c we have so much video data, we have this thing that can project the future in 2d pixel space. Projecting into the future in 3d world space is actually what the endgame for robotics is and I imagine depending on how complex that 3d world model is, a working model for projecting into 3d space could be waaaaaay smaller. It's just that the equivalent data is not as easily available on…

There are also models that are trained to generate 3D models from a picture. Use it on videos, and also train it on output generated by video games.

Re: Video generation models as world simulators

#77

I like that this one shows some "fails", and not just the top of the top results: For example, the surfer is surfing in the air at the end: https://cdn.openai.com/tmp/s/prompting_7.mp4 Or this "breaking" glass that does not break, but spills liquid in some weird way: https://cdn.openai.com/tmp/s/discussion_0.mp4 Or the way this person walks: https://cdn.openai.com/tmp/s/a-woman-wearing-a-green-dress-a... Or wherever…

> Or wherever this map is coming from:

> https://cdn.openai.com/tmp/s/a-woman-wearing-purple-overalls...

Notice also that that a roughly 6 seconds there is a third hand putting the map away.

Re: Video generation models as world simulators

#78
post #69

Damn, even minecraft videos being simulated, this is crazy to see from OpenAI. Edit, changed the links to the direct ones! https://cdn.openai.com/tmp/s/simulation_6.mp4 https://cdn.openai.com/tmp/s/simulation_7.mp4

As someone who's played probably too many hours of minecraft, these videos are nauseating. The way that all of the individual pieces exist, but have no consistency is terrifying. Random FoV changes, switching apparent texture packs, raytracing on or off, it's all still switching back and forth from moment to moment. These videos honestly give me less confidence in the approach, simply because I don't know that the mo…

As someone who is a human born in the 80s, this shit wasn't supposed to even be theoretically possible.

People didn't think this was possible even a year ago.

"Gives me less confidence"

Cmon man...

Re: Video generation models as world simulators

#79
Video will be especially important for language models to grasp physical actions that are instinctive and obvious to humans but not explicitly detailed in text or video captions. I mentioned this in 2022:

https://twitter.com/LechMazur/status/1607929403421462528

https://twitter.com/LechMazur/status/1619032477951213568

Post reply on HN