Live data from Hacker News

SANA-WM, a 2.6B open-source world model for 1-minute 720p video

nvlabs.github.io

11–20 of 162 posts

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#14

First video with the guy walking the mountain in snow has consistency issues with the cave entrance. Which is "expected" at this model size?!

Most videos seem to have some issues like that, e.g. the book on the table in the library video takes up different shapes every now and then.

The 'Refiner' effect seems to do the opposite if the examples are representative as in all cases the 1-st stage images look better than the 'refined' ones. Less clutter, more realistic, less 'cowbell' for those who know the phrase.

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#16
I struggle with these world models from the perspective of video games (so this post is a particular perspective).

I'm not a game developer myself, but some of my favorite games carry a deep sense of intentionality. For instance, there is typically not a single item misplaced in a FromSoftware game (or, for instance, Lies of P -- more recently). Almost every object is placed intentionally.

Games which lack this intentionality often feel dead in contrast. You run into experiences which break immersion, or pull you out of the experience that the developer is trying to convey to you.

It's difficult for me to imagine world models getting to a place where this sort of intentionality is captured. The best frontier LLMs fail to do this in writing (all the time), and even in code, and the surface of experiences for those mediums often feel "smaller" than the user interaction profile of a video game.

It's not clear how these world models could be used modularly by humans hoping to develop intentional experiences? I don't know much about their usage (LLMs are somewhat modular: they can produce text, humans can work on it, other LLMs can work on it). Is the same true for the video output here?

All this to say, I'm impressed with these world models, but similar to LLMs with writing, it's not really clear what it is that we are building towards? We are able to create less satisfying, less humane experiences faster? Perhaps the most immediate benefit is the ability for robotic systems to simulate actions (by conjuring a world, and imagining the implications).

In general, I have the feeling that we are hurtling towards a world with less intentionality behind all the things we experience. Everything becomes impersonal, more noisy, etc.

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#18

What’s the long term utility of world models? There’s no doubt they’re technically impressive, but what does one do with it?

They can be base models for a bunch of things. Turning text-conditioned video generation models into robotics VLAs is a fun exercise.

This one is probably too small to be useful for that, and not diverse enough? But I could be wrong.

Post reply on HN