Live data from Hacker News

SANA-WM, a 2.6B open-source world model for 1-minute 720p video

nvlabs.github.io

101–110 of 162 posts

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#101
post #84

silly question: what's "world" about what's being generated here? is the an actual abstract representation of physical space (like, eg, a game-engine style scene graph?) or does it just mean "this video generator is more coherent physically than other video generators"

World in this context means that these videos are interactive, just like a video game. In the linked examples you can see the keyboard and mouse inputs. The model is trained to maintain about a minute of scene consistency so you can look around and objects out of view will reappear when you look back in that direction.

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#103

warning: viewing the videos that auto play on that page shot up my downloads to 350Mbps on that page

I only noticed after more than an hour with the page left open in a tab. Is it really streaming and re-streaming the same videos? There's too much to cache so it keeps re-transferring them indefinitely? I hope nobody leaves that page open on a metered or capped network connection. I'm surprised github hasn't suspended the page. Are AI researchers so used to burning through compute and network resources that they don'…

They don't even notice it happening, it is not a conscious thought not to fix it.

Empathizing about problems you don't face is a hard product/ux and management skill. Facebook famously simulated 2G on Tuesdays 10 years ago[1] for example to get their employees to see the problems their users have.[2]

People don't to put effort in noticing(solving comes next) problems they don't face. It is why things like a11y and i18n need regulation like ADA etc.

[1] https://engineering.fb.com/2015/10/27/networking-traffic/bui...

[2]While it would be hard to attribute directly, GraphQL and to an extent React probably was influenced by these kind of things

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#104
post #16

I struggle with these world models from the perspective of video games (so this post is a particular perspective). I'm not a game developer myself, but some of my favorite games carry a deep sense of intentionality. For instance, there is typically not a single item misplaced in a FromSoftware game (or, for instance, Lies of P -- more recently). Almost every object is placed intentionally. Games which lack this inten…

i've played multiple AAA(+) games before AI 'was a thing' that have had textures/elements, like bulletin boards or posters, where even on cursory glances (not zooming in or ADS) you can easily see literally "Lorem Ipsum" instead of lore or story which would have helped build atmosphere

LLMs had nothing to do with this

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#105
post #16

I struggle with these world models from the perspective of video games (so this post is a particular perspective). I'm not a game developer myself, but some of my favorite games carry a deep sense of intentionality. For instance, there is typically not a single item misplaced in a FromSoftware game (or, for instance, Lies of P -- more recently). Almost every object is placed intentionally. Games which lack this inten…

I suspect these models will be like old Gutenberg's printing press. A rapid rise in the amount of content; most of it not that great. However the sheer volume will result in even more high quality content actually being created in aggregate. Put another way, the average game quality will go down, but the actual rate of "Great" games will go up.

But these aren't great games. They are not even good. They are just tech demos with nothing of interest to gamers.

Why do I need more slopware? I have an entire Steam library of excellent games that deserve to be played first.

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#106
post #16

I struggle with these world models from the perspective of video games (so this post is a particular perspective). I'm not a game developer myself, but some of my favorite games carry a deep sense of intentionality. For instance, there is typically not a single item misplaced in a FromSoftware game (or, for instance, Lies of P -- more recently). Almost every object is placed intentionally. Games which lack this inten…

Consider instead the possibility this may be used as a rendering layer for data backing it. Instead of shipping three-dimensional models and GBs of textures, you can ship a couple photos or a blueprint file or and a detailed text description for significantly less storage. Now imagine the world model can adapt the styling of this world on the fly, where every person‘s experience could be unique in terms of visuals, but consistent in terms of the gameplay.

It’s been my belief for several years that this is how the future of games will be constructed. Data in the background, game engine for rules application/ physics execution/orchestration/maybe low-poly rendering, an AI world model taking low resolution input in generating customized visuals/effects/textures/everything, even camera location, but still constrained by concrete rules in the game engine.

I’m certain one day it might all be handled by AI, but the above seems much more realistic and achievable that expecting AI to do all of these things, at one, correctly, every frame.

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#107

Model weights coming "soon" == currently vaporware. So the weights aren't even open, how can this be "open-source"? Everyone is right to be skeptical of this coming from a 2.8B model. Weights or it didn't happen.

To be fair, their whole codebase is open-source, which is better than most open-weight models. But I do agree with the sentiment.

https://github.com/NVlabs/Sana

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#108
post #16

I struggle with these world models from the perspective of video games (so this post is a particular perspective). I'm not a game developer myself, but some of my favorite games carry a deep sense of intentionality. For instance, there is typically not a single item misplaced in a FromSoftware game (or, for instance, Lies of P -- more recently). Almost every object is placed intentionally. Games which lack this inten…

> Games which lack this intentionality often feel dead in contrast Like for instance... Dwarf Fortress? Minecraft? Generative AI is just another method to go procedural generation. Not necessarily a better way. Or you could even argue that procedural generation is a form of generative AI... But either way, there are games where the lack of intentionality is central to the appeal.

minecraft itself is a blank slate until the player or modder or whatever puts all that intentionality in.

its a very dead game on its own. they are still very intentional about adding and changing the tools by which you make your own fun though.

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#109
post #16

I struggle with these world models from the perspective of video games (so this post is a particular perspective). I'm not a game developer myself, but some of my favorite games carry a deep sense of intentionality. For instance, there is typically not a single item misplaced in a FromSoftware game (or, for instance, Lies of P -- more recently). Almost every object is placed intentionally. Games which lack this inten…

That is interesting. And it's an AI critique I haven't heard before. Would you consider it possible that the way non-intentionally placed items break the game immersion for you is because they appear in such a way that you think you can interact with them in a certain way, but you can't? Like if there's an extra door in the house you're trying to get into, but that door doesn't really open, then in your mind that bre…

its reasonably the same "reversion to the mean" or "not x, but y"

the intentionally placed tree serves no particular in-game job mechanically. it instead points your eyes to the right place when you walk up the path, and then again when you look back down from above.

when they're saying everything is intentionally placed, they mean everything, whether it looks important or not. It's all directed to a cohesive core

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#110
post #84

silly question: what's "world" about what's being generated here? is the an actual abstract representation of physical space (like, eg, a game-engine style scene graph?) or does it just mean "this video generator is more coherent physically than other video generators"

A world-model is one that predicts the next state of a simulated world given the current state and optionally some action from an agent inhabiting the world. It is quite analogous to a language-model that predicts the next word.

That world-state can be anything, but in the last year or two, the term has taken a narrower meaning: a video generation model that reacts naturally to game-like controls, as if it was simulating a videogame. But there's no additional state behind the video frames.

Post reply on HN