Live data from Hacker News

Video generation models as world simulators

openai.com

51–60 of 171 posts

Re: Video generation models as world simulators

#51

I think people might be missing what this enables. It can make plausible continuations of video, with realistic physics. What happens if this gets fast enough to work _in real time_. Connect this to a robot that has a real time camera feed. Have it constantly generate potential future continuations of the feed that it's getting -- maybe more than one. You have an autonomous robot building a real time model of the wor…

You're talking about an agent with a world model used for planning. Actually generating realistic images is not really needed as the world model operates in its own compressed abstraction.

Check out V-Jepa for such a system: https://ai.meta.com/blog/v-jepa-yann-lecun-ai-model-video-jo...

Re: Video generation models as world simulators

#52

I think people might be missing what this enables. It can make plausible continuations of video, with realistic physics. What happens if this gets fast enough to work _in real time_. Connect this to a robot that has a real time camera feed. Have it constantly generate potential future continuations of the feed that it's getting -- maybe more than one. You have an autonomous robot building a real time model of the wor…

> Connect this to a robot that has a real time camera feed. Have it constantly generate potential future continuations of the feed that it's getting -- maybe more than one. You have an autonomous robot building a real time model of the world around it and predicting the future. Give it some error correction based on well each prediction models the actual outcome and I think you're _really_ close to AGI. In theory, ye…

I’ve been dying for someone to make a Civilization AI.

It might not be too crazy of an idea - would love to see a model fine-tuned on sequences of moves.

The biggest limitation of video game AI currently is not theory, but hardware. Once home compute doubles a few more times, we’ll all be running GPT-4 locally and a competent Civilization AI starts to look realistic.

Re: Video generation models as world simulators

#53

I think people might be missing what this enables. It can make plausible continuations of video, with realistic physics. What happens if this gets fast enough to work _in real time_. Connect this to a robot that has a real time camera feed. Have it constantly generate potential future continuations of the feed that it's getting -- maybe more than one. You have an autonomous robot building a real time model of the wor…

how would you define AGI?

Re: Video generation models as world simulators

#54

Earlier quoted context omitted.

i mean everyone's mind builds a convincing internal simulation of reality and it's so good that most people think they're directly experiencing reality.

So what happens to someone suffering a psychotic episode, their reality gets distorted? But what reality though if it’s all an internal simulation? I think there’s partly an internal simulation to some aspect of reality but there’s a lot more to it.

The world model is not the world. It's the old map and territory thing.

Re: Video generation models as world simulators

#55

I think people might be missing what this enables. It can make plausible continuations of video, with realistic physics. What happens if this gets fast enough to work _in real time_. Connect this to a robot that has a real time camera feed. Have it constantly generate potential future continuations of the feed that it's getting -- maybe more than one. You have an autonomous robot building a real time model of the wor…

Imagine putting on some AR goggles staring at a painting in a Museum Then immediately jumping into an entire VR world based off the painting generated by an AI rendering it out on the fly

BlockadeLabs has been doing a 3D text to skybox and not exactly runtime at the moment but I have seen it work in a headset and it definitely feels like the future.

Re: Video generation models as world simulators

#56

I think people might be missing what this enables. It can make plausible continuations of video, with realistic physics. What happens if this gets fast enough to work _in real time_. Connect this to a robot that has a real time camera feed. Have it constantly generate potential future continuations of the feed that it's getting -- maybe more than one. You have an autonomous robot building a real time model of the wor…

I totally agree that a system like Sora is needed. By itself, it’s insufficient. With a multimodal model that can reason properly, then we get AGI or rather ASI (artificial super intelligence) due to many advantages over humans such as context length, access to additional sensory modalities (infrared, electroreception, etc), much broader expertise, huge bandwidth, etc.

future successor to Sora + likely successor to GPT-4 = ASI

See my other comment here: https://news.ycombinator.com/item?id=39391971

Re: Video generation models as world simulators

#57

Earlier quoted context omitted.

> Connect this to a robot that has a real time camera feed. Have it constantly generate potential future continuations of the feed that it's getting -- maybe more than one. You have an autonomous robot building a real time model of the world around it and predicting the future. Give it some error correction based on well each prediction models the actual outcome and I think you're _really_ close to AGI. In theory, ye…

I’ve been dying for someone to make a Civilization AI. It might not be too crazy of an idea - would love to see a model fine-tuned on sequences of moves. The biggest limitation of video game AI currently is not theory, but hardware. Once home compute doubles a few more times, we’ll all be running GPT-4 locally and a competent Civilization AI starts to look realistic.

I think it's also a matter of "shape". Like, GPT4 solves one "shape" of problem, given tokens, predict the next token. That's all it does, that's the only problem it has to solve.

A Civilization AI would have many problem "shapes". What do I research? Where do I build my city, what buildings do I build, how do I move my units, what units do I build, what improvements do I build, when do I declare war, what trade deals do I accept, etc, etc. Each of those is fundamentally different, and you can maybe come up with a scheme to make them all into the same "shape", but then that ends up being harder to train. I would be interested to see a good solution to this problem.

Re: Video generation models as world simulators

#58

I like that this one shows some "fails", and not just the top of the top results: For example, the surfer is surfing in the air at the end: https://cdn.openai.com/tmp/s/prompting_7.mp4 Or this "breaking" glass that does not break, but spills liquid in some weird way: https://cdn.openai.com/tmp/s/discussion_0.mp4 Or the way this person walks: https://cdn.openai.com/tmp/s/a-woman-wearing-a-green-dress-a... Or wherever…

> For example, the surfer is surfing in the air at the end

Maybe it’s been watching snowboarding videos and doesn’t quite understand the difference.

Re: Video generation models as world simulators

#59
post #34

While the Sora videos are impressive, are these really world simulators? While some notion of real-world physics probably exists somewhere within the model, doesn’t all the completely artificial training data corrupt it? Reasoning, logic, formal systems, and physics exist in a seemingly completely different, mathematical space than pure video. This is just a contrived, interesting viewpoint of the technology, right?

> Reasoning, logic, formal systems, and physics exist in a seemingly completely different, mathematical space than pure video.

That's not true, AI systems in general have pretty strong mathematical proofs going back decades on what they can theoretically do, the problem is compute and general feasibility. AIXItl in theory would be able to learn reasoning, logic, formal systems, physics, human emotions, and a great deal of everything else just from watching videos. They would have to be videos of varied and useful things, but even if they were not, you'd at least get basic reasoning, logic, and physics.

Re: Video generation models as world simulators

#60

I like that this one shows some "fails", and not just the top of the top results: For example, the surfer is surfing in the air at the end: https://cdn.openai.com/tmp/s/prompting_7.mp4 Or this "breaking" glass that does not break, but spills liquid in some weird way: https://cdn.openai.com/tmp/s/discussion_0.mp4 Or the way this person walks: https://cdn.openai.com/tmp/s/a-woman-wearing-a-green-dress-a... Or wherever…

I've also noticed on some of the featured videos that there are some perspective/parallax errors. The human subjects in some are either oversized compared to background people, or they end up on horizontal planes that don't line up properly. It's actually a bit vertigo-inducing! It is still very remarkable
Post reply on HN