Live data from Hacker News

Video generation models as world simulators

openai.com

91–100 of 171 posts

Re: Video generation models as world simulators

#91
post #8

If there's one thing I've always wanted, it's shitty video knockoffs of real life. Can't wait to stream some AI hallucinations.

As if what you consume normally is actual real life. With blue screens, VFX, etc. you are already watching knockoffs of real life, and the shitty will become indistinguishable from reality before long.

Re: Video generation models as world simulators

#92
The video with the two MTBs going downhill: it seems to me that the long left turn that begins a few second into the video is way too long. It's easy to misjudge that kind of things (try to draw a road race track by looking at a single lap of it) but it could end up below the point where it started, or too close to it to be physically realistic. I was expecting to see a right turn at any moment but it kept going left. It could be another consequence of the lack of real knowledge about the world, similar to the glass shattering example at the end of the article.

Re: Video generation models as world simulators

#93

Earlier quoted context omitted.

Sounds like simulation theory is closer and closer to being proven.

Except there is always an original at the root. There’s no way to prove that’s not us.

Haven't you heard of "turtles all the way down"?

Re: Video generation models as world simulators

#94
If they would allow this (maybe a premium+ model) they could soon destroy the whole porn industry. not the websites, but the (often abused) sex workers. Everyone could describe that fetish they are into and get it visualized instantly without the need of physical human suffering to produce these videos.

I know its a delicate topic people (especially in the US) don't want to speak about at all, but damn, this is a giant market and could do humanity good if done well.

Re: Video generation models as world simulators

#95

I am a newbie to this area. Honest questions: Is this generating videos as streaming content e.g. like a mp4 video. As far as I can see, it is doing that. Is it possible for AI to actually produce the 3d models? What kind of compute resources are required to produce the 3d models.

You can feed the output to NeRF or Gaussian Splat generators to produce 3d models:

https://twitter.com/BenMildenhall/status/1758224827788468722

https://twitter.com/ScottieFoxTTV/status/1758272455603327455

The key is that the video has spatial consistency. Once you've got that, then other existing tech can take the output and infer actual spatial forms.

Re: Video generation models as world simulators

#96

Earlier quoted context omitted.

I’ve been dying for someone to make a Civilization AI. It might not be too crazy of an idea - would love to see a model fine-tuned on sequences of moves. The biggest limitation of video game AI currently is not theory, but hardware. Once home compute doubles a few more times, we’ll all be running GPT-4 locally and a competent Civilization AI starts to look realistic.

I am 100% certain that the training of such an AI will result in winning a game without ever building a single city* and 1,000 other exploits before being nerfbatted enough to play a 'real' game. (That doesn't mean I don't want to see the ridiculousness it comes up with!) * https://www.youtube.com/watch?v=6CZEEvZqJC0

I knew it, I knew it! It would be a Spiffing Brit video.

That guy is a genius at finding exploits in computer games. I don't know how he does it, I think you need to play a fair bit of each game before you find these little corners of the ruleset.

Re: Video generation models as world simulators

#97

I think people might be missing what this enables. It can make plausible continuations of video, with realistic physics. What happens if this gets fast enough to work _in real time_. Connect this to a robot that has a real time camera feed. Have it constantly generate potential future continuations of the feed that it's getting -- maybe more than one. You have an autonomous robot building a real time model of the wor…

What I find interesting is that b/c we have so much video data, we have this thing that can project the future in 2d pixel space. Projecting into the future in 3d world space is actually what the endgame for robotics is and I imagine depending on how complex that 3d world model is, a working model for projecting into 3d space could be waaaaaay smaller. It's just that the equivalent data is not as easily available on…

imagine it going a few dimensions further, what will happen when i tell this person 'this'. how will this affect the social graph and my world state :)

Re: Video generation models as world simulators

#99

Maybe this says more about me than about the technology, but I found the consistency of the Minecraft simulation super impressive.

I was wondering how feasible it would be to make a Minecraft agent that had a running feed of the past few seconds, continued it off w/ SORA, fed the continuation into a (relatively) simple policy translator that just pulled out what the video showed as player inputs, and the inputted that.

Presumably, this would work for non-minecraft applications, but Minecraft has a really standardized interface layer

Re: Video generation models as world simulators

#100
post #2

I find it wild that this model does not have explicit 3D prior, yet learns to generate videos with such 3D consistency, you can directly train a 3D representation (NeRF-like) from those videos: https://twitter.com/BenMildenhall/status/1758224827788468722

You aren't looking carefully enough. I find so many inconsistencies in these examples. Perspectives that are completely wrong when the camera rotates. Windows that shift perspective, patios that are suddenly deep/shallow. Shadows that appear/disappear as the camera shifts. In other examples; paths, objects, people suddenly appearing or disappearing out of nowhere. A stone turning into a person. A horse that suddenly has a second head, then becomes a separate horse with only two legs.

It is impressive at a glance, but if you pay attention, it is more like dreaming than realism (images conjured out of other images, without attention to long term temporal, spatial, and causal consistency). I'm hardly more impressed that Google's deep dream, which is 10 years old.

Post reply on HN