Live data from Hacker News

SANA-WM, a 2.6B open-source world model for 1-minute 720p video

nvlabs.github.io

131–140 of 162 posts

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#131
post #127

Gist: > 720p, 1-min video generation with 6-DoF camera control As nl said, > The model is out here: https://huggingface.co/Efficient-Large-Model/SANA-Video_2B_7... README says "intended for research use only" Code license is Apache 2.0 Model license (nvidia open...) says Models are commercially usable. You are free to create and distribute Derivative Models (As usual: model output is unrestricted, and also unprotecta…

The linked model does not claim to support camera control, it doesn't appear to be SANA-WM. You might be getting confused.

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#132
post #16

I struggle with these world models from the perspective of video games (so this post is a particular perspective). I'm not a game developer myself, but some of my favorite games carry a deep sense of intentionality. For instance, there is typically not a single item misplaced in a FromSoftware game (or, for instance, Lies of P -- more recently). Almost every object is placed intentionally. Games which lack this inten…

By and large I agree, but it doesn’t need to be either/or. Many of the most popular games in the past decade are procedurally generated and have nothing “intentionally” placed (apart from tuning/tweaking the balance of the seeding algorithms).

I've had good luck with using LLMs to create procedural content engines for my game prototypes. So the distinction between AI and procedural might get even blurrier.

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#133
post #16

I struggle with these world models from the perspective of video games (so this post is a particular perspective). I'm not a game developer myself, but some of my favorite games carry a deep sense of intentionality. For instance, there is typically not a single item misplaced in a FromSoftware game (or, for instance, Lies of P -- more recently). Almost every object is placed intentionally. Games which lack this inten…

One thing is robotics. Both for training robotics AI, and to let robots test hypothetical actions before comitting to them. I don't think world models are stable enough for either yet The other is creating multi-modal models with a better understanding of our world. LLMs often fail at incredibly basic spatial reasoning ("someone left a package in front of your apartment, describe going there", or the "should I drive…

We are seeing increasing evidence that these sort of video world models are terrible options for useful rollouts of the physical dynamics of the environment. It is hypothesized that you can get them to be better than simulators by training on physics simulation data, but then the question becomes, why not use the simulator directly?

There are a lot of areas where predictive models make sense in the robotics stack, but doing it with "video world models" as is trendy this year is likely a bet in the wrong direction according to the evidence we have been amassing in the last 6 months.

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#134
post #16

I struggle with these world models from the perspective of video games (so this post is a particular perspective). I'm not a game developer myself, but some of my favorite games carry a deep sense of intentionality. For instance, there is typically not a single item misplaced in a FromSoftware game (or, for instance, Lies of P -- more recently). Almost every object is placed intentionally. Games which lack this inten…

What does intentionality mean in the context of a world model generated game-world? I guess true human intention would have been throw out the window already at that point. One aspect of intentionality is that there’ll be a narrative payoff when you find something you find interesting. In videogames, the world is mostly pre-designed, so the designer has to predict what you’ll be interested in for the most part (In pe…

> there’ll be a narrative payoff

Fromsoft is perfectly happy for you to miss all of the direct exposition. It's as they intended and most people do. The intentionality of their world still draws people in and gives the world a sense of groundedness that keeps people coming back and separates it from the pale imitations. It's more than them being good at 'Vibes'.

The environment is built on the bones of a greater ongoing narrative that is intentionally obscured, even from the player who reads everything.

Dark Souls is a world in a constant cycle of Rebirth, Decay, the struggle against Entropy. All civilizations, at the end of the series, are stacked one upon another in an endless expanse of ash and dust as you bear witness to an permanent eclipse, a fading star, as time itself dissolves and the last fire fades.

Before that heavy handed stuff though, the simple matter of the direction your character travels reinforces these motifs. Down to the deepest depths and you'll witness what remains of the first civilizations. Climb up and you see the desperate attempts by the powerful to impose a false order that they hoped could forestall the inevitable.

You can even shatter the illusion of a golden order in the first game if you find the extremely missable secret boss. It couldn't be any more clearly 'said' if you were interested in paying attention.

Adding an AI model to explain or 'improv' the story of the world would destroy the whole purpose.

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#135

Outputting video of that quality/consistency at 1 minute, for a 2.6B model seems insane?

It's because it is insane/misleading. It's a two stage process, scroll to the key features:

> A dedicated 17B long-video refiner sharpens texture, motion, and late-window quality on top of the long-rollout backbone.

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#136

Earlier quoted context omitted.

This intentionality in the application of AI is very confusing for folks because at first glance it seems like it should just work. It seems to, even. Whereas if you hand a router to someone with a flush trim but in it and ask them to clean up the edge of a table they will take one look at it and nope away from that dangerous spinning thing. If they have the mind to give it a shot and despite a quality tool and bit t…

Just like with mass-produced materials vs hand-crafted stuff, you're gonna have a lot of crap quality and rare, expensive good quality stuff.

And most people can’t just spin up a furniture factory at their whim and call themselves a designer. AI gives everybody with the slightest gumption a fully-functional, “initially plausible crap” factory at their fingertips, so everybody with actual skills gets lost in a sea of useless garbage.

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#137

Earlier quoted context omitted.

But these aren't great games. They are not even good. They are just tech demos with nothing of interest to gamers. Why do I need more slopware? I have an entire Steam library of excellent games that deserve to be played first.

Agreed, these aren't even games currently. I am just saying that world models will lower the barrier to entry to making games. Which might mean that 1 in 1000 of the lower-barrier-to-entry-people might someday makes a great game. So more great games in aggregate, but more bad games on average.

And theoretically AI does a great job at helping HR filter unqualified candidates, and it helps candidates optimize their resumes and application strategies to help them land the right role. So people should be landing dream roles left-and-right. Is that how it’s working?

In reality, I don’t see any of this trending towards the theoretical happy path everybody always talks about. Most people give up trying to find something good on Amazon and just buy whatever vaguely plausible knock-off garbage shows up in the first few search results. Most people just take any job interview they’re offered even if it sucks. Most HR people don’t use it to enhance the quality of their decisions — it replaces their decision-making roles in many respects.

I’m an art school graduate and talk am in many art discussion communities. This is causing a massive industry-wide morale crater. In any sort of art, it damn near eliminates the reward of craftsmanship in favor of marketing useless trend-of-the-week bullshit. Far fewer people enter a market that can’t sustain them. The idea that this is going to create ‘more artists’ and therefore that must mean there must be more skilled artists is fantasy. The skills you learn by prompting are not even on the same track to learning how to create things yourself. You essentially become a high-school intern acting as an art director, commissioning pieces. It’s instant gratification for people who don’t care enough about something to learn how to do it for real.

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#138

warning: viewing the videos that auto play on that page shot up my downloads to 350Mbps on that page

I only noticed after more than an hour with the page left open in a tab. Is it really streaming and re-streaming the same videos? There's too much to cache so it keeps re-transferring them indefinitely? I hope nobody leaves that page open on a metered or capped network connection. I'm surprised github hasn't suspended the page. Are AI researchers so used to burning through compute and network resources that they don'…

> Are AI researchers so used to burning through compute and network resources that they don't stop to think about a webpage that will autoplay and loop multiple HD videos?

I’m sure they’ll give their Claude instance a stern talking-to.

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#139
post #127

Gist: > 720p, 1-min video generation with 6-DoF camera control As nl said, > The model is out here: https://huggingface.co/Efficient-Large-Model/SANA-Video_2B_7... README says "intended for research use only" Code license is Apache 2.0 Model license (nvidia open...) says Models are commercially usable. You are free to create and distribute Derivative Models (As usual: model output is unrestricted, and also unprotecta…

[deleted]

Re: SANA-WM, a 2.6B open-source world model for 1-minute 720p video

#140

Earlier quoted context omitted.

Agreed, these aren't even games currently. I am just saying that world models will lower the barrier to entry to making games. Which might mean that 1 in 1000 of the lower-barrier-to-entry-people might someday makes a great game. So more great games in aggregate, but more bad games on average.

And theoretically AI does a great job at helping HR filter unqualified candidates, and it helps candidates optimize their resumes and application strategies to help them land the right role. So people should be landing dream roles left-and-right. Is that how it’s working? In reality, I don’t see any of this trending towards the theoretical happy path everybody always talks about. Most people give up trying to find so…

Just to be clear, I'm visualizing the usage of world models that can consistently render visual and interactive renderings of a specified world. I think interacting with them will be markedly different than interacting with many text based LLMs (though I don't know, I have never had direct access to one).

I don't think these will create "artist" in any sense, but I do think it will lower the barrier dramatically for people creating games. Most people will interact with it like Lieutenant Barclay interacting with the holodeck, doing little more than wish fulfillment. But I think a few people will be able to interact with it in ways that create art.

In no way am I implying that the net net of AI will be good for humanity as a whole (I think that is too big a question), but I do think the power of World Models will probably result in a far more people being able to say "I have created a game".

I honestly don't have anything useful to say about what LLMs are doing to many human fields. I can understand how frustrating it must feel to see LLMs demonstrate superhuman "skill" (I don't really think they are skilled) at orders of magnitude less cost than a good artist. It isn't just that they don't seem to innovate (only permute), it is that they will literally take even the tiniest bit of creativity and novelty and immediately fine tune and create derivative works on any idea at scale. I can see how that might really demotivate any desire to push the boundaries of art for any human being.

Post reply on HN