Live data from Hacker News

Diffusion models are real-time game engines

gamengen.github.io

371–380 of 430 posts

Re: Diffusion models are real-time game engines

#371

Earlier quoted context omitted.

>It has no notion of game state (so you can kill an enemy, turn your back, then turn around again) Well you see a wall you turn around then turn back the wall is still there. With enough training data the model will be able to pick up the state of the enemy because it has ALREADY learned the state of the wall due to much more numerous data on the wall. It's probably impractical to do this, but this is only a stepping…

> Well you see a wall you turn around then turn back the wall is still there. With enough training data the model will be able to pick up the state of the enemy because it has ALREADY learned the state of the wall due to much more numerous data on the wall. It's really important to understand that ALL THE MODEL KNOWS is a mapping of [pixels, input] -> new pixels. It has zero knowledge of game state. The wall is still…

>It's really important to understand that ALL THE MODEL KNOWS is a mapping of [pixels, input] -> new pixels. It has zero knowledge of game state.

This is false. What occurs in inside the model is unknown. It arranges pixel input and produces pixel output as if it actually understands game state. Like LLMs we don't actually fully understand what's going on internally. You can't assume that models don't "understand" things just because the high level training methodology only includes pixel input and output.

>The only "state" that is known is the last few frames of the game screen. Because of this, it's simply not possible for the game model to know if an enemy should be shown as dead or alive once it has been off-screen for longer than those few frames. It also means that if you keeping turning away and towards an enemy, it could teleport around. Once it's off the screen for those few frames, the model will have forgotten about it.

This is true. But then one could say it knows game state for up to a few frames. That's different from saying the model ONLY knows pixel input and pixel output. Very different.

There are other tricks for long term memory storage as well. Think Radar. Radar will capture the state of the enemy beyond just visual frames so the model won't forget an enemy was behind them.

Game state can also be encoded into some frame pixels at the bottom lines. The Model can pick up on these associations.

edit: someone mentioned that the game state lasts past a few frames.

>If you're trying to make a new game, then you need new frames to train the model on.

Right so for a generative model you would instead of training the model on one game you would train it on multitudes of games. The model would then based off of a seed number output a new type of game.

Alternatively you could have a model generate a model.

All of what I'm saying is of course speculative. As I said, this model is a stepping stone for the future. Just like the LLM which is only trivially helpful now, the LLM can be a stepping stone for replacing programmers all together.

Re: Diffusion models are real-time game engines

#372

Earlier quoted context omitted.

What would the model provide if not what we see on the screen?

The environment and everything in it. “Everything” would mean all objects and the elements they’re made of, their rules on how they interact and decay. A modularized ecosystem i guess, comprised of “sub-systems” of sorts. The other model, that provides all interaction (cause for effect) could either be run artificially or be used interactively by a human - opening up the possibility for being a tree : ) This all woul…

What you’re asking for doesn’t make sense.

Re: Diffusion models are real-time game engines

#374

After some discussion in this thread, I found it worth pointing out that this paper is NOT describing a system which receives real-time user input and adjusts its output accordingly, but, to me, the way the abstract is worded heavily implied this was occurring. It's trained on a large set of data in which agents played DOOM and video samples are given to users for evaluation, but users are not feeding inputs into the…

I also thought this, but refer back to the paper, not the abstract: > A is the set of key presses and mouse movements… > …to condition on actions, we simply learn an embedding A_emb for each action So, it’s clear that in this model the diffusion process is conditioned by embedding A that is derived from user actions rather than words. Then a noised start frame is encoded into latents and concatenated on to the noise…

The agent never interacts with the simulator during training or evaluation. There is no user, there is only an agent which trained to play the real game and which produced the sequences of game frames and actions that were used to train the simulator and to provide ground truth sequences of game experience for evaluation. Their evaluation metrics are all based on running short simulations in the diffusion model which are initiated with some number of conditioning frames taken from the real game engine. Statements in the paper like: "GameNGen shows that an architecture and model weights exist such that a neural model can effectively run a complex game (DOOM) interactively on existing hardware." are wildly misleading.

Re: Diffusion models are real-time game engines

#375

After some discussion in this thread, I found it worth pointing out that this paper is NOT describing a system which receives real-time user input and adjusts its output accordingly, but, to me, the way the abstract is worded heavily implied this was occurring. It's trained on a large set of data in which agents played DOOM and video samples are given to users for evaluation, but users are not feeding inputs into the…

You are incorrect, this is an interactive simulation that is playable by humans. > Figure 1: a human player is playing DOOM on GameNGen at 20 FPS. The abstract is ambiguously worded which has caused a lot of confusion here, but the paper is unmistakably clear about this point. Kind of disappointing to see this misinformation upvoted so highly on a forum full of tech experts.

If the generative model/simulator can run at 20FPS, then obviously in principle a human could play the game in simulation at 20 FPS. However, they do no evaluation of human play in the paper. My guess is that they limited human evals to watching short clips of play in the real engine vs the simulator (which conditions on some number of initial frames from the engine when starting each clip...) since the actual "playability" is not great.

Re: Diffusion models are real-time game engines

#376

After some discussion in this thread, I found it worth pointing out that this paper is NOT describing a system which receives real-time user input and adjusts its output accordingly, but, to me, the way the abstract is worded heavily implied this was occurring. It's trained on a large set of data in which agents played DOOM and video samples are given to users for evaluation, but users are not feeding inputs into the…

Ehhh okay, I'm not as convinced as I was earlier. Sorry for misleading. There's been a lot of back-and-forth. I would've really liked to see a section of the paper explicitly call out that they used humans in real time. There's a lot of sentences that led me to believe otherwise. It's clear that they used a bunch of agents to simulate gameplay where those agents submitted user inputs to affect the gameplay and they c…

It's funny how academic writing works. Authors rarely produce many unclear or ambiguous statements where the most likely interpretation undersells their work...

Re: Diffusion models are real-time game engines

#377

Earlier quoted context omitted.

Okay, I think you're right. My mistake. I read through the paper more closely and I found the abstract to be a bit misleading compared to the contents. Sorry.

Don't worry. The paper is not very well written.

Academic authors are consistently better at editing away unclear and ambiguous statements which make their work seem less impressive compared to ones which make their work seem more impressive. Maybe it's just a coincidence, lol.

Re: Diffusion models are real-time game engines

#378
post #29

Earlier quoted context omitted.

That feels like the endgame of video game generation. You select an art style, a video and the type of game you'd like to play. The game is then generated in real-time responding to each action with respect to the existing rule engine. I imagine a game like that could get so convincing in its details and immersiveness that one could forget they're playing a game.

IIRC, both 2001 (1968) and Solaris (1972) depict that kind of things as part of alien euthanasia process, not as happy endings

Well, 2001 is actually a happy ending, as Dave is reborn as a cosmic being. Solaris, at least in the book, is an attempt by the sentient ocean to communicate with researchers through mimics.

Re: Diffusion models are real-time game engines

#379

After some discussion in this thread, I found it worth pointing out that this paper is NOT describing a system which receives real-time user input and adjusts its output accordingly, but, to me, the way the abstract is worded heavily implied this was occurring. It's trained on a large set of data in which agents played DOOM and video samples are given to users for evaluation, but users are not feeding inputs into the…

We can't assess the quality of gameplay ourselves of course (since the model wasn't released), but one author said "It's playable, the videos on our project page are actual game play." (https://x.com/shlomifruchter/status/1828850796840268009) and the video on top of https://gamengen.github.io/ starts out with "these are real-time recordings of people playing the game". Based on those claims, it seems likely that they did get a playable system in front of humans by the end of the project (though perhaps not by the time the draft was uploaded to arXiv).

Re: Diffusion models are real-time game engines

#380

Earlier quoted context omitted.

Then why do monsters become blurry smudgy messes when shot? That looks like a video compression artifact of a neural network attempting to replicate low-structure image (source material contains guts exploding, very un-structured visual).

Uh, maybe because monster death animations make up a small part of the training material (ie. gameplay) so the model has not learned to reproduce them very well? There cannot be "video compression artifacts" because it hasn’t even seen any compressed video during training, as far as I can see. Seriously, how is this even a discussion? The article is clear that the novel thing is that this is real-time frame generatio…

In a sense, poorly reproducing rare content is a form of compression artifact. Ie, since this content occurs rarely in the training set, it will have less impact on the gradients and thus less impact on the final form of the model. Roughly speaking, the model is allocating fewer bits to this content, by storing less information about this content in its parameters, compared to content which it sees more often during training. I think this isn't too different from certain aspects of images, videos, music, etc., being distorted in different ways based on how a particular codec allocates its available bits.
Post reply on HN