Live data from Hacker News

Packing Input Frame Context in Next-Frame Prediction Models for Video Generation

lllyasviel.github.io

21–29 of 29 posts

Re: Packing Input Frame Context in Next-Frame Prediction Models for Video Generation

#23
post #10

Earlier quoted context omitted.

I haven't bothered with video gen because I'm too impatient but isn't Wan pretty good too on regular hardware?

Wan 2.1 (and Hunyuan and LTXV, in descending ordee of overall video quality but each has unique strengths) work well—but slow, except LTXV—for short (single digit seconds at their usual frame rates — 16 for WAN, 24 for LXTV, I forget for Hunyuan) videos on consumer hardware. But this blows them entirely out of the water on the length it can handle, so if it does so with coherence and quality across general prompts (e…

For completeness, I should note I'm talking about the 14B i2v and t2v WAN 2.1 models; there are others in the family, notably a set of 1.3B models that are presumably much faster, but I haven't worked with them as much

Re: Packing Input Frame Context in Next-Frame Prediction Models for Video Generation

#24
post #4

Funny how it really wants people to dance. Even the guy sitting down for an interview just starts dancing sitting down.

It's a peculiar and fascinating observation you make.

With static images, we always look for eyes.

With video, we always look for dancing.

Re: Packing Input Frame Context in Next-Frame Prediction Models for Video Generation

#26
post #15
post #4

Funny how it really wants people to dance. Even the guy sitting down for an interview just starts dancing sitting down.

Presumably they're dancing because it's in the prompt. You could change the prompt to have them do something else (but that would be less fun!)

I'm no expert but are you sure there is a prompt?

Re: Packing Input Frame Context in Next-Frame Prediction Models for Video Generation

#27
post #15

Earlier quoted context omitted.

Presumably they're dancing because it's in the prompt. You could change the prompt to have them do something else (but that would be less fun!)

I'm no expert but are you sure there is a prompt?

Yes, while the page here does not directly mention the prompts, the linked paper does, and the linked code repo shows that prompts are used as well.

Re: Packing Input Frame Context in Next-Frame Prediction Models for Video Generation

#28

Earlier quoted context omitted.

I'm no expert but are you sure there is a prompt?

Yes, while the page here does not directly mention the prompts, the linked paper does, and the linked code repo shows that prompts are used as well.

100%. I don't think I've ever even come across an I2V model that didn't require at least a positive prompt. Some people get around it by integrating a vision LLM into their ComfyUI workflows however.

Re: Packing Input Frame Context in Next-Frame Prediction Models for Video Generation

#29

Earlier quoted context omitted.

I'm no expert but are you sure there is a prompt?

Yes, while the page here does not directly mention the prompts, the linked paper does, and the linked code repo shows that prompts are used as well.

Ah yeah you're right - they seem to just really like giving dancing prompts. I guess they work well due to the training set.
Post reply on HN