Live data from Hacker News

Video generation models as world simulators

openai.com

141–150 of 171 posts

Re: Video generation models as world simulators

#141
post #104

Earlier quoted context omitted.

> Or the way this person walks: > https://cdn.openai.com/tmp/s/a-woman-wearing-a-green-dress-a ... Also, why does she have a umbrella sticking out from her lower back?

I suppose the lady usually has an umbrella in this kind of situation, so it felt the umbrella should be included in some way: https://youtu.be/492tGcBP568

In truth, that's not a woman in a green dress, it's a bunch of penguins disguised as a woman in a green dress. That explains the gait. As to the umbrella, they assumed that humans, intelligent as we are, always carry polar bear protection around.

Re: Video generation models as world simulators

#142
post #2

I find it wild that this model does not have explicit 3D prior, yet learns to generate videos with such 3D consistency, you can directly train a 3D representation (NeRF-like) from those videos: https://twitter.com/BenMildenhall/status/1758224827788468722

You aren't looking carefully enough. I find so many inconsistencies in these examples. Perspectives that are completely wrong when the camera rotates. Windows that shift perspective, patios that are suddenly deep/shallow. Shadows that appear/disappear as the camera shifts. In other examples; paths, objects, people suddenly appearing or disappearing out of nowhere. A stone turning into a person. A horse that suddenly…

You can literally run 3D algorithms like NeRF or COLMAP on those videos (check the tweet I sent), it's not my opinion, those videos are sufficiently 3D consistent that you can extract 3D geometry from them

Surely it's not perfect, but this was not the case for previous video generation algorithms

Re: Video generation models as world simulators

#143
> We empirically find that training on videos at their native aspect ratios improves composition and framing. We compare Sora against a version of our model that crops all training videos to be square, which is common practice when training generative models. The model trained on square crops (left) sometimes generates videos where the subject is only partially in view. In comparison, videos from Sora (right)s have improved framing.

Every cv preprocessing pipeline is in shambles now.

Re: Video generation models as world simulators

#144

Earlier quoted context omitted.

Where do you find the last two?

Part of this website changes after the video finished and switches to the next video. There is no way to control it. These are both "X wearing Y taking a pleasant stroll in Z during W"

nvm you can control it, when you click on the variables, there is a dropdown to select values for the variables.

Re: Video generation models as world simulators

#145

Earlier quoted context omitted.

That is creepy...

Take a look at these adorable kangaroos to relax: https://cdn.openai.com/tmp/s/an-adorable-kangaroo-wearing-bl... https://cdn.openai.com/tmp/s/an-adorable-kangaroo-wearing-a-...

I think you forgot the most adorable one /s

https://cdn.openai.com/tmp/s/an-adorable-kangaroo-wearing-bl...

no but for real this seems to be the most "adorable" one:

https://cdn.openai.com/tmp/s/an-adorable-kangaroo-wearing-bl...

Re: Video generation models as world simulators

#146

I think people might be missing what this enables. It can make plausible continuations of video, with realistic physics. What happens if this gets fast enough to work _in real time_. Connect this to a robot that has a real time camera feed. Have it constantly generate potential future continuations of the feed that it's getting -- maybe more than one. You have an autonomous robot building a real time model of the wor…

Figure out how to incorporate a quantum computer as a prediction engine in this idea, and you've got quite the robot on your hands. :)

(and throw this in for good measure https://www.wired.com/story/this-lab-grown-skin-could-revolu... heh)

Re: Video generation models as world simulators

#147
This is a totally silly thought, but I still want to get it out there.

> Other interactions, like eating food, do not always yield correct changes in object state

Can this be because we just don't shoot a lot of people eating? I think it is general advice to not show people eating on camera for various reasons. I wonder if we know if that kind of topic bias exists in the dataset.

Re: Video generation models as world simulators

#148

Earlier quoted context omitted.

One more French brain we didn't manage to keep. The drain is just crazy at this point.

Do you have a list ?

I don't, I just seem to have this moment of "oh, him as well" regularly.

And I get it, I went to the valley as well for some times, the money is better, the taxes are lower, you get more opportunities, meet more talented people and projects are way cooler.

Re: Video generation models as world simulators

#150
post #8

If there's one thing I've always wanted, it's shitty video knockoffs of real life. Can't wait to stream some AI hallucinations.

So is the only thing you watch on TV, unedited documentary footage?

Also - ironic choice of username considering this comment!

Post reply on HN