Live data from Hacker News

Packing Input Frame Context in Next-Frame Prediction Models for Video Generation

lllyasviel.github.io

11–20 of 29 posts

Re: Packing Input Frame Context in Next-Frame Prediction Models for Video Generation

#13
post #10
post #3

This guy is a genius; for those who don’t know he also brought us ControlNet. This is the first decent video generation model that runs on consumer hardware. Big deal and I expect ControlNet pose support soon too.

I haven't bothered with video gen because I'm too impatient but isn't Wan pretty good too on regular hardware?

LTX-Video isn't quite the same quality as Wan, but the new distilled 0.9.6 version is pretty good and screamingly fast.

https://github.com/Lightricks/LTX-Video

Re: Packing Input Frame Context in Next-Frame Prediction Models for Video Generation

#14
post #10
post #3

This guy is a genius; for those who don’t know he also brought us ControlNet. This is the first decent video generation model that runs on consumer hardware. Big deal and I expect ControlNet pose support soon too.

I haven't bothered with video gen because I'm too impatient but isn't Wan pretty good too on regular hardware?

Wan 2.1 is solid but you start to get pretty bad continuity / drift issues when genning more than 81 frames (approx 5 seconds of video) whereas FramePack lets you generate 1+ minute.

Re: Packing Input Frame Context in Next-Frame Prediction Models for Video Generation

#15
post #4

Funny how it really wants people to dance. Even the guy sitting down for an interview just starts dancing sitting down.

Presumably they're dancing because it's in the prompt. You could change the prompt to have them do something else (but that would be less fun!)

Re: Packing Input Frame Context in Next-Frame Prediction Models for Video Generation

#16
post #12
post #5

looks like the only motion it can do...is to dance

It can dance if it wants to... It can leave LLMs behind... 'Cause LLMs don't dance, and if they don't dance, well, they're no friends of mine.

That's a certified bop! ;) You should get elybeatmaker to do a remix!

Edit: I didn't realize that this was actually a reference to Men Without Hats - The Safety Dance. I was referencing a different parody/allusion to that song!

Re: Packing Input Frame Context in Next-Frame Prediction Models for Video Generation

#17
post #12
post #5

looks like the only motion it can do...is to dance

It can dance if it wants to... It can leave LLMs behind... 'Cause LLMs don't dance, and if they don't dance, well, they're no friends of mine.

The AI Safety dance?

Re: Packing Input Frame Context in Next-Frame Prediction Models for Video Generation

#18
post #3

This guy is a genius; for those who don’t know he also brought us ControlNet. This is the first decent video generation model that runs on consumer hardware. Big deal and I expect ControlNet pose support soon too.

He also brought us IC-Light! I wonder why he's still contributing to open source... Surely all the big companies have made him huge offers. He's so talented

Re: Packing Input Frame Context in Next-Frame Prediction Models for Video Generation

#19
post #10
post #3

This guy is a genius; for those who don’t know he also brought us ControlNet. This is the first decent video generation model that runs on consumer hardware. Big deal and I expect ControlNet pose support soon too.

I haven't bothered with video gen because I'm too impatient but isn't Wan pretty good too on regular hardware?

Wan 2.1 (and Hunyuan and LTXV, in descending ordee of overall video quality but each has unique strengths) work well—but slow, except LTXV—for short (single digit seconds at their usual frame rates — 16 for WAN, 24 for LXTV, I forget for Hunyuan) videos on consumer hardware. But this blows them entirely out of the water on the length it can handle, so if it does so with coherence and quality across general prompts (especially if it is competitive with WAN and Hunyuan on trainability for concepts it may not handle normally) it is potentially a radical game changer.

Re: Packing Input Frame Context in Next-Frame Prediction Models for Video Generation

#20
post #3

This guy is a genius; for those who don’t know he also brought us ControlNet. This is the first decent video generation model that runs on consumer hardware. Big deal and I expect ControlNet pose support soon too.

He also brought us IC-Light! I wonder why he's still contributing to open source... Surely all the big companies have made him huge offers. He's so talented

I think he is working on his Ph.D. at Stanford. I assume whatever offers he has haven't been attractive enough to abandon that, whether he’ll still be doing open work or get sucked into the bowels of some proprietary corporate behemoth afterwards remains to be seen, but I suspect he won't have trouble monetizing his skills either way.
Post reply on HN