Live data from Hacker News

Video generation models as world simulators

openai.com

161–170 of 171 posts

Re: Video generation models as world simulators

#161

Earlier quoted context omitted.

This argument always bothers me. The probability of being in any simulation is conditional on the one above, which necessarily decreases exponentially. Any simulation running an equivalent simulation will do so much slower, so you get a geometric series of degrading probabilities. The rate of decay will be massive, imagine how long and how much resources it would take us to simulate our universe, even in a hand waved…

That's not a good argument, because we have no way of knowing what is the ration between our time and that of the universe where our simulation runs. Even if it takes hours to generate one second of our universe, we only experience our own time. Besides, time is not absolute, and having it run slower near massive objects or when objects accelerate would be a neat trick to save on compute power needed to simulate a un…

I wouldn’t have used the slower argument but rather the information encoding argument. A simulation must necessarily encode less information than the universe being simulated, in fact substantially less. It wouldn’t be possible to encode the same information as the root universe in any nested simulation, even the first level. That would require the root entire universe to encode. This information encoding problem is what geometrically gets worse as you nest. At a certain point the simulation must be so simplified and lossy of simulated information that it’s got no meaningful information and the simulation isn’t representative of anything.

However - just because probabilistic reasoning explodes and the likelihood of something is vanishing doesn’t make it true.

First, the prior assumption is the universe can be simulated in any meaningful way at any substantial scale. That’s not at all obvious that the ability to simulate is high enough fidelity to lead to the complexity we see around us without some higher dimensional universe simulating what we see and the realities we see are achieved through dimensional reduction and absurdly powerful technology. This is also a probability in the conditional probability and I would not put it at 1.0. I would actually make it quite small, but its term as a prior will be significant.

Second, the prior is that whatever root universe that exists has yet to achieve the simulation in the flow of time, assuming time started at some discrete point. Our observations lead us to conclude time and space both emerged at a discrete point. The coalescence of the modern universe, evolution of life itself, emergence of intelligent beings, the technology required to simulate an enormous highly complex universe in its entirety, etc, are all priors. These are non trivial factors to consider and greatly reduce the likelihood of the simulation theory.

Third, it’s possible the clock rate of the simulation is fast enough that the simulation operates much faster than time evolves in the root universe, but to the original posts point, without enormous lossy optimizations, the nested universes can’t run at a faster clock rate in their simulation than the first level simulation. This is partially related to the information encoding problem but not directly. I don’t agree it geometrically gets worse, but it doesn’t get better either without further greatly reducing the quality of the simulation. That means either the quality converges to zero very fast, or they run at a synchronicity of the first level universe, requiring 1:1 time. Assuming it actually simulates the universe and not just some sort of occlusion scoped to you as an individual, that might mean it’ll take billions of years within the first level simulation for each layer of the nesting. This seems practically unlikely even in a simulated universe, so either those layers must not achieve a nesting or they must converge to simulations that have lost so much fidelity they simulate nothing very quickly.

Re: Video generation models as world simulators

#162

Earlier quoted context omitted.

That's not a good argument, because we have no way of knowing what is the ration between our time and that of the universe where our simulation runs. Even if it takes hours to generate one second of our universe, we only experience our own time. Besides, time is not absolute, and having it run slower near massive objects or when objects accelerate would be a neat trick to save on compute power needed to simulate a un…

I wouldn’t have used the slower argument but rather the information encoding argument. A simulation must necessarily encode less information than the universe being simulated, in fact substantially less. It wouldn’t be possible to encode the same information as the root universe in any nested simulation, even the first level. That would require the root entire universe to encode. This information encoding problem is…

Well first you imply a base universe is finite. That is not a given at all.

You don't need to simulate the full universe. Just the experience of consciousness inside it. You don't even have to simulate full consciousness for every 'conscious' being. In fact, I've always seen the simulation argument as a thought experiment arguing for consciousness being more fundamental than matter. There is no need to imagine a human made computer simulating an entire universe in subatomic detail for this thought experiment to intrigue us.

We being able to pinpoint a start of all time is actually a pretty good argument for it being simulated. Why would we be able to calculate a 'start time' for reality? That is not obvious to be a necessity at a base universe at all. There are theoretical cosmologies out there that do away with that need to conceptualize a universe.

The simulated universe doesn't have to run time faster then 'real time' at the base universe at all. In fact, running slower would be a feature if the beings in the base universe wished to escape into the simulation for whatever reason.

Re: Video generation models as world simulators

#163

Earlier quoted context omitted.

I am 100% certain that the training of such an AI will result in winning a game without ever building a single city* and 1,000 other exploits before being nerfbatted enough to play a 'real' game. (That doesn't mean I don't want to see the ridiculousness it comes up with!) * https://www.youtube.com/watch?v=6CZEEvZqJC0

But wouldn't this be amazing for the developer to fix a lot of edge cases/bugs?

Maybe, maybe not. The stochastic, black-box nature of the current wave of ML systems gives me a gut feeling that using them like this is more of a Monkey's Paw wish granter than useful tool without a lot of refinement first. Time will tell!

Re: Video generation models as world simulators

#164
post #107

Ugh, AI generated images everywhere is already annoying enough. Now we're gonna have these factitious videos clogging up everything, and I'll have to explain my old neighbor that Biden did infact not eat a fetus again and again.

100%. It's actually gotten even more dull once they started fixing the fingers. But it's too much; you start to realise that it's just so uninspired. Maybe what this will ultimately do is allow good writers to bring their ideas to life (I hope).

Re: Video generation models as world simulators

#165

I think people might be missing what this enables. It can make plausible continuations of video, with realistic physics. What happens if this gets fast enough to work _in real time_. Connect this to a robot that has a real time camera feed. Have it constantly generate potential future continuations of the feed that it's getting -- maybe more than one. You have an autonomous robot building a real time model of the wor…

Adding to this: Sora was most likely trained on video that's more like what you'd normally see on YouTube or in a clip art or media licensing company collection. Basically, video designed to look good as a part of a film or similar production.

So right now, Sora is predicting "Hollywood style" content, with cuts, camera motions, etc... all much like what you'd expect to see in an edited film.

Nothing stops someone (including OpenAI) from training the same architecture with "real world captures".

Imagine telling a bunch of warehouse workers that for "safety" they all need to wear a GoPro-like action camera on their helmets that record everything inside the work area. Run that in a bunch of warehouses with varying sizes, content, and forklifts, and then pump all of that through this architecture to train it. Include the instructions given to the staff from the ERP system as well as the transcribed audio as the text prompt.

Ta-da.

You have yourself an AI that can control a robot using the same action camera as its vision input. It will be able to follow instructions from the ERP, listen to spoken instructions, and even respond with a natural voice. It'll even be able to handle scenarios such as spills, breaks, or other accidents... just like the humans in its training data did. This is basically what vehicle auto-pilots do, but on steroids.

Sure, the computer power required for this is outrageously expensive right now, but give it ten to twenty years and... no more manual labour.

Re: Video generation models as world simulators

#166

Earlier quoted context omitted.

I am 100% certain that the training of such an AI will result in winning a game without ever building a single city* and 1,000 other exploits before being nerfbatted enough to play a 'real' game. (That doesn't mean I don't want to see the ridiculousness it comes up with!) * https://www.youtube.com/watch?v=6CZEEvZqJC0

I knew it, I knew it! It would be a Spiffing Brit video. That guy is a genius at finding exploits in computer games. I don't know how he does it, I think you need to play a fair bit of each game before you find these little corners of the ruleset.

Idk maybe he uses some sort of fuzzer

Re: Video generation models as world simulators

#167

Earlier quoted context omitted.

Really interesting how this goes against my intuition. I would have imagined that it's infinitely easier to analyze a camera stream of the real world, then generate a polygonal representation of what you see (like you would do for a videogame) and then make AI decisions for that geometry. Instead the way that AI is going they rather skip it all and work directly on pixel data. Understanding of 3d geometry, perspectiv…

> then generate a polygonal representation of what you see It's really not that surprising since, to be honest, meshes suck. They're pretty general graphs but to actually work nicely they have to have really specific topological characteristics. Half of the work you do with meshes is repeatedly coaxing them back into a sane topology after editing them.

Do we have anything better than meshes that is as generally useful though?

Re: Video generation models as world simulators

#169
post #115

Earlier quoted context omitted.

One more French brain we didn't manage to keep. The drain is just crazy at this point.

What would fix the issue?

Based on talking to my coworkers from Europe on the West Coast (some have more nuanced position, but some were outright "everyone in tech in their right mind moves away from Europe"), nothing short-term.

If you forget specifics, and consider on the abstract level what the differences are... Let's say there was an equal pile of "resources" per person available in Europe and the US. The way this pile is (abstractly) distributed in Europe is egalitarian and safety-net focused; in the US it is distributed more unequally, closer to some imperfect approximation of merit. Most of the (real) advantages and disadvantages that people bring up for the US stem from that. The more of this approximation of merit you have, the more "resources" you'd have in US. No matter what the specific slopes are, unless one place is much richer (might be the US anyway), at some point these lines cross. The higher the person is above this point the more it makes sense for them to go to the US...

There are also 2nd order effects like other people above that point having already gone (not just from Europe, from everywhere in the world), making the US more attractive, probably. That might matter more for top talent.

And although this probably doesn't matter for the top talent, "regular" Europeans can actually have the cake and eat it too - make the money in the US, then (in old age or if something happens) move back home and avail themselves of the welfare state. A non-German guy who worked in Germany for a few years told me that's what he'd do if he was German - working in Germany sucks, but being lazy in Germany is wonderful, so he'd move to the US then move back ;)

Re: Video generation models as world simulators

#170

Earlier quoted context omitted.

I wouldn’t have used the slower argument but rather the information encoding argument. A simulation must necessarily encode less information than the universe being simulated, in fact substantially less. It wouldn’t be possible to encode the same information as the root universe in any nested simulation, even the first level. That would require the root entire universe to encode. This information encoding problem is…

Well first you imply a base universe is finite. That is not a given at all. You don't need to simulate the full universe. Just the experience of consciousness inside it. You don't even have to simulate full consciousness for every 'conscious' being. In fact, I've always seen the simulation argument as a thought experiment arguing for consciousness being more fundamental than matter. There is no need to imagine a huma…

I didn’t imply the base universe is finite just it has to be orders of magnitude larger than the simulation.

Yes but if you only simulate a consciousness there is no nesting, otherwise on what does the consciousness run on? That’s the whole idea behind as you lose fidelity with reality the simulations becomes less capable until it’s unable to simulate anything.

Note nesting isn’t necessary for the probability argument to be true just many simulations. But the vanishingly small likelihood of not being a simulation depends on the nesting idea.

Again, yes, it’s possible what we observe is a simple simulation of experience but I don’t think this is the “almost certainly a simulation” argument.

Post reply on HN