Live data from Hacker News

Stable Video Diffusion

stability.ai

21–30 of 316 posts

Re: Stable Video Diffusion

#22
I admit I'm ignorant about these model's inner workings, but I don't understand why text is the chosen input format for these models.

It was the same for image generation, where one needed to produce text prompts to create the image, and stuff like img2img and Controlnet that allowed things like controlling poses and inpainting, or having multiple prompts with masks controlling which part of the image is influenced by which prompt.

Re: Stable Video Diffusion

#23

In the video towards the bottom of the page, there are two birds (blue jays), but in the background there are two identical buildings (which look a lot like the CN Tower). CN Tower is the main landmark of Toronto, whose baseball team happens to be the Blue Jays. It's located near the main sportsball stadium downtown. I vaguely understand how text-to-image works, and so it makes sense that the vector space for "blue j…

I don’t spend a lot of time keeping up with the space, but I could have sworn I’ve seen a demo that allowed you to iterate in the way you’re suggesting. Maybe someone else can link it.

Re: Stable Video Diffusion

#24

In the video towards the bottom of the page, there are two birds (blue jays), but in the background there are two identical buildings (which look a lot like the CN Tower). CN Tower is the main landmark of Toronto, whose baseball team happens to be the Blue Jays. It's located near the main sportsball stadium downtown. I vaguely understand how text-to-image works, and so it makes sense that the vector space for "blue j…

Assuming we can post links, you mean this video: https://youtu.be/G7mihAy691g?si=o2KCmR2Uh_97UQ0N

Also, maybe you can't edit post facto, but when you give prompts, would you not be able to say : two blue jays but no CN tower

Re: Stable Video Diffusion

#25
It makes sense that they had to take out all of the cuts and fades from the training data to improve results.

I’m the background section of the research paper they mention “temporal convolution layers”, can anyone explain what that is? What sort of training data is the input to represent temporal states between images that make up a video? Or does that mean something else?

Re: Stable Video Diffusion

#26
post #14

Can this be used for porn?

Very unusual comment. I do not think so as the chance of constructing a fleshy eldritch horror is quite high.

How is that not the first question to ask? Porn has proven to be a fantastic litmus test of fast market penetration when it comes to new technologies.

Re: Stable Video Diffusion

#27

In the video towards the bottom of the page, there are two birds (blue jays), but in the background there are two identical buildings (which look a lot like the CN Tower). CN Tower is the main landmark of Toronto, whose baseball team happens to be the Blue Jays. It's located near the main sportsball stadium downtown. I vaguely understand how text-to-image works, and so it makes sense that the vector space for "blue j…

Not exactly what you're asking for, but AnimateDiff has introduced creating gifs to SD. Still takes quite a bit of tweaking IME.

Re: Stable Video Diffusion

#28

In the video towards the bottom of the page, there are two birds (blue jays), but in the background there are two identical buildings (which look a lot like the CN Tower). CN Tower is the main landmark of Toronto, whose baseball team happens to be the Blue Jays. It's located near the main sportsball stadium downtown. I vaguely understand how text-to-image works, and so it makes sense that the vector space for "blue j…

Assuming we can post links, you mean this video: https://youtu.be/G7mihAy691g?si=o2KCmR2Uh_97UQ0N Also, maybe you can't edit post facto, but when you give prompts, would you not be able to say : two blue jays but no CN tower

Yes, its called a negative prompt. Idk if txt2video has it, but both llms and stable-diffusion have it so I'd assume its good to go.

Re: Stable Video Diffusion

#29
A seemingly off topic question, but with enough compute and optimization, could you eventually simulate “reality”?

Like, at this point, what are the technical counters to the assertion that our world is a simulation?

Post reply on HN