Live data from Hacker News

SANA-WM, a 2.6B open-source world model for 1-minute 720p video

nvlabs.github.io

91–100 of 162 posts

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#93
post #43

First video with the guy walking the mountain in snow has consistency issues with the cave entrance. Which is "expected" at this model size?!

Remember the first Will Smith spaghetti?

Yeah it got ridiculed and people wrote it off as that was somehow the limit and that wasn't going to change which seems to be the common premise to which people launch their criticisms of AI

And those same people forget that its been 3 years since that awful will smith spaghetti video to what we have today which is the beginning of controllable real time videos aka games

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#94

Earlier quoted context omitted.

Games. Build campaigns in hours instead of months. Make it possible for users to create their own campaigns, move the action to different game worlds - 'gimme Mario Kart in the ${favourite_game} world', etc.

Yeah, but is this really that great? Are these models going to remember the town you wandered through on your session yesterday and want to return to? Imagine playing Read Dead Redemption 2 and you attempt to ride your horse from Saint Denis to Valentine and Valentine no longer exists, or is a completely different town located half a mile off from where it was originally. I just don't see how this would work...

Remember code generation ? 6 years ago you could barely get it to generate anything complex.

Remember video generation? 3 years ago the will smith spaghetti video came out.

You see how this trend will only continue? Game development is going to get really weird.

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#95

All video models are terrible at consistency. Even closed source ones. Seedance 2.0, Kling 3 are regarded the best closed source video models we have. I have subscribed to a few AI video subreddits, consensus atm is they are good for anything but long form videos with humans. No surprises that we're very good at spotting even the most subtle differences while looking at other people.

Relax its only been 3 years, it's going to get a lot better not worse from here on.

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#97

So, where is the download? I can't find it on Github, and on your web page the download button is disabled. Also, will this run on RTX 4090 with 24GB memory? Thank you!

There's a 5 second version available: https://huggingface.co/Efficient-Large-Model/SANA-Video_2B_7...

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#98
post #71

I tried watching the cave video and I was immediately overcome with nausea. I've never experienced anything like that before in my life. Wild. I can't say I'm looking forward to an AI video future.

When I installed very high quality (CRI 98, R9 94, virtually flicker free) light bulbs to one of my apartment rooms I had headaches and felt occasional confusion for about a week while being in that room, so I had to slowly increase the amount of time the light bulbs were turned on for. To my understanding my brain was very used to the way objects and lighting looked in that particular room so it needed to rewire some knowledge given that I've spent many thousands of hours in that room with previous light bulbs.

I'm curious if a younger me would have adapted much faster.

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#99
post #16

I struggle with these world models from the perspective of video games (so this post is a particular perspective). I'm not a game developer myself, but some of my favorite games carry a deep sense of intentionality. For instance, there is typically not a single item misplaced in a FromSoftware game (or, for instance, Lies of P -- more recently). Almost every object is placed intentionally. Games which lack this inten…

I suspect these models will be like old Gutenberg's printing press. A rapid rise in the amount of content; most of it not that great. However the sheer volume will result in even more high quality content actually being created in aggregate.

Put another way, the average game quality will go down, but the actual rate of "Great" games will go up.

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#100
post #16

I struggle with these world models from the perspective of video games (so this post is a particular perspective). I'm not a game developer myself, but some of my favorite games carry a deep sense of intentionality. For instance, there is typically not a single item misplaced in a FromSoftware game (or, for instance, Lies of P -- more recently). Almost every object is placed intentionally. Games which lack this inten…

That is interesting. And it's an AI critique I haven't heard before.

Would you consider it possible that the way non-intentionally placed items break the game immersion for you is because they appear in such a way that you think you can interact with them in a certain way, but you can't?

Like if there's an extra door in the house you're trying to get into, but that door doesn't really open, then in your mind that breaks the integrity of the game's systems. If so, I think the LLM response is that there are no more doors that don't open and that the world can be generated as needed.

No computer can handle the complexity of even a small town. But it would be possible, at least in the future, to generate the part of the world you interact with, which would heighten the emersion.

Post reply on HN