Earlier quoted context omitted.
Controlnet is adapted to video today, the issues are that it's very slow. Haven't you seen the insane quality of videos on civitai?
I have seen them, the workflows to create those videos are extremely labor intensive. Control net lets you maintain poses between frames, it doesn’t solve the temporal consistency of small details.
Stable Video Diffusion
51–60 of 316 posts
Re: Stable Video Diffusion
#52In the video towards the bottom of the page, there are two birds (blue jays), but in the background there are two identical buildings (which look a lot like the CN Tower). CN Tower is the main landmark of Toronto, whose baseball team happens to be the Blue Jays. It's located near the main sportsball stadium downtown. I vaguely understand how text-to-image works, and so it makes sense that the vector space for "blue j…
I feel like we're close too, but for another reason.
For although I love SD and these video examples are great... It's a flawed method: they never get lighting correctly and there are many incoherent things just about everywhere. Any 3D artist or photographer can immediately spot that.
However I'm willing to bet that we'll soon have something much better: you'll describe something and you'll get a full 3D scene, with 3D models, source of lights set up, etc.
And the scene shall be sent into Blender and you'll click on a button and have an actual rendering made by Blender, with correct lighting.
Wanna move that bicycle? Move it in the 3D scene exactly where you want.
That is coming.
And for audio it's the same: why generate an audio file when soon models shall be able to generate the various tracks, with all the instruments and whatnots, allowing to create the audio file?
That is coming too.
Re: Stable Video Diffusion
#53A seemingly off topic question, but with enough compute and optimization, could you eventually simulate “reality”? Like, at this point, what are the technical counters to the assertion that our world is a simulation?
Re: Stable Video Diffusion
#54In the video towards the bottom of the page, there are two birds (blue jays), but in the background there are two identical buildings (which look a lot like the CN Tower). CN Tower is the main landmark of Toronto, whose baseball team happens to be the Blue Jays. It's located near the main sportsball stadium downtown. I vaguely understand how text-to-image works, and so it makes sense that the vector space for "blue j…
I don’t spend a lot of time keeping up with the space, but I could have sworn I’ve seen a demo that allowed you to iterate in the way you’re suggesting. Maybe someone else can link it.
Re: Stable Video Diffusion
#55Looks like I'm still good for my bet with some friends that before 2028 a team of 5-10 people will create a blockbuster style movie that today costs 100+ million USD on a shoestring budget and we won't be able to tell.
The first full-length AI generated movie will be an important milestone for sure, and will probably become a "required watch" for future AI history classes. I wonder what the Rotten Tomatoes page will look like.
Re: Stable Video Diffusion
#56Instance One : Act as a top tier Hollywood scenarist, use the public available data for emotional sentiment to generate a storyline, apply the well known archetypes from proven blockbusters for character development. Move to instance two.
Instance Two: Act as top tier producer. {insert generated prompt}. Move to instance three.
Instance Three: Generate Meta-humans and load personality traits. Move to instance four.
Instance Four: Act as a top tier director.{insert generated prompt}. Move to instance five.
Instance Five: Act as a top tier editor.{insert generated prompt}. Move to instance six.
Instance Six: Act as a top tier marketing and advertisement agency.{insert generated prompt}. Move to instance seven.
Instance Seven: Act as a top tier accountant, generate an interface to real-time ROI data and give me the results on an optimized timeline into my AI induced dream.
Personal GPT: Buy some stocks, diversify my portfolio, stock up on synthetic meat, bug-coke and Soma. Call my mom and tell her I made it.
Re: Stable Video Diffusion
#57In the video towards the bottom of the page, there are two birds (blue jays), but in the background there are two identical buildings (which look a lot like the CN Tower). CN Tower is the main landmark of Toronto, whose baseball team happens to be the Blue Jays. It's located near the main sportsball stadium downtown. I vaguely understand how text-to-image works, and so it makes sense that the vector space for "blue j…
I don’t spend a lot of time keeping up with the space, but I could have sworn I’ve seen a demo that allowed you to iterate in the way you’re suggesting. Maybe someone else can link it.
Re: Stable Video Diffusion
#58Earlier quoted context omitted.
Very unusual comment. I do not think so as the chance of constructing a fleshy eldritch horror is quite high.
How is that not the first question to ask? Porn has proven to be a fantastic litmus test of fast market penetration when it comes to new technologies.
Re: Stable Video Diffusion
#59Earlier quoted context omitted.
I have seen them, the workflows to create those videos are extremely labor intensive. Control net lets you maintain poses between frames, it doesn’t solve the temporal consistency of small details.
People use animatediff’s motion module (or other models that have cross frame attention layers). Consistency is close to being solved.
Re: Stable Video Diffusion
#60Fascinating leap forward. It makes me think of the difference between ancestral and non-ancestral samplers, e.g. Euler vs Euler Ancestral. With Euler, the output is somewhat deterministic and doesn't vary with increasing sampling steps, but with Ancestral, noise is added to each step which creates more variety but is more random/stochastic. I assume to create video, the sampler needs to lean heavily on the previous f…
So this example was posted an hour ago, and it's jumping all over the place frame to frame (somewhat weak temporal consistency). The author appears to have used pretty straight-forward text2img + Animatediff:
https://www.reddit.com/r/StableDiffusion/comments/180no09/on...
Fixing that frame to frame jitter related to animation is probably the most in-demand thing around Stable Diffusion right now.
Animatediff motion painting made a splash the other day:
https://www.reddit.com/r/StableDiffusion/comments/17xnqn7/ro...
It's definitely an exciting time around SD + animation. You can see how close it is to reaching the next level of generation.