In the video towards the bottom of the page, there are two birds (blue jays), but in the background there are two identical buildings (which look a lot like the CN Tower). CN Tower is the main landmark of Toronto, whose baseball team happens to be the Blue Jays. It's located near the main sportsball stadium downtown. I vaguely understand how text-to-image works, and so it makes sense that the vector space for "blue j…
> Has anyone come across a solution where model can iterate (eg, with prompts like "move the bicycle to the left side of the photo")? It feels like we're close. I feel like we're close too, but for another reason. For although I love SD and these video examples are great... It's a flawed method: they never get lighting correctly and there are many incoherent things just about everywhere. Any 3D artist or photographer…
Stable Video Diffusion
81–90 of 316 posts
Re: Stable Video Diffusion
#82Looks like I'm still good for my bet with some friends that before 2028 a team of 5-10 people will create a blockbuster style movie that today costs 100+ million USD on a shoestring budget and we won't be able to tell.
Back in the mid 90s to 2010 or so, graphical improvements were hailed as photorealistic only to be improved upon with each subsequent blockbuster game.
I think we're in a similar phase with AI[0]: every new release in $category is better, gets hailed as super fantastic world changing, is improved upon in the subsequent Two Minute Papers video on $category, and the cycle repeats.
[0] all of them: LLMs, image generators, cars, robots, voice recognition and synthesis, scientific research, …
Re: Stable Video Diffusion
#83To be clear, this remains wicked, wicked, wicked exciting.
Re: Stable Video Diffusion
#84In the video towards the bottom of the page, there are two birds (blue jays), but in the background there are two identical buildings (which look a lot like the CN Tower). CN Tower is the main landmark of Toronto, whose baseball team happens to be the Blue Jays. It's located near the main sportsball stadium downtown. I vaguely understand how text-to-image works, and so it makes sense that the vector space for "blue j…
This is not the flex you think it is. You don't have to like sports, but snarking on people who do doesn't make you intellectual, it just makes you come across as a douchebag, no different than a sports fan making fun of "D&D nerds" or something.
Re: Stable Video Diffusion
#85I'm still puzzled as to how these "non-commercial" model licenses are supposed to be enforceable. Software licenses govern the redistribution of the software , not products produced with it. An image isn't GPL'd because it was produced with GIMP.
Visual Studio Community (and many other products) only allows "non-commercial" usage. Sounds like it limits what you can do with what you produce with it. At the end of the day, a license is a legal contract. If you agree that an image which you produce with some software will be GPL'ed, it's enforceable. As an example, see the Creative Commons license, ShareAlike clause: > If you remix, transform, or build upon the…
you can put whatever you want in a contract, doesn't mean it's enforceable
Re: Stable Video Diffusion
#86Re: Stable Video Diffusion
#87Earlier quoted context omitted.
It's really not. Don't get me wrong, this is insanely cool, but it's still a long way from good enough to be truly disruptive.
One year. All of Hollywood falls.
If anything, a deluge of cheap AI generated movies is going to lead to a flight to quality. The big studios will be more powerful because they will reap the productivity gains and use traditional techniques to smooth out the rough edges.
Re: Stable Video Diffusion
#88Earlier quoted context omitted.
People use animatediff’s motion module (or other models that have cross frame attention layers). Consistency is close to being solved.
Temporal consistency is improving, but “close to being solved” is very optimistic.
And that’s just for diffusion focused approaches like ours. There are probably other techniques from the token flow or nerf family of approaches close to breakout levels of quality, tons of talented researchers working on that too.
Re: Stable Video Diffusion
#89Earlier quoted context omitted.
One year. All of Hollywood falls.
No offense, but this is absolutely delusional. As long as people can "clock" content generated from these models, it will be treated by consumers as low-effort drivel, no matter how much actual artistic effort goes in the exercise. Only once these systems push through the threshold of being indistinguishable from artistry will all hell break loose, and we are still very far from that. Paint-by-numbers low-effort mark…
CGI in films used to be obvious all the time no matter how good the artists using it, now it's everywhere and only noticeable when that's the point; the gap from Tron to Fellowship of the Ring was 19.5 years.
My guess is the analogy here puts the quality of existing genAI somewhere near the equivalent of early TV CGI, given its use in one of the Marvel title sequences etc., but it is just an analogy and there's no guarantees of anything either way.
Re: Stable Video Diffusion
#90In the video towards the bottom of the page, there are two birds (blue jays), but in the background there are two identical buildings (which look a lot like the CN Tower). CN Tower is the main landmark of Toronto, whose baseball team happens to be the Blue Jays. It's located near the main sportsball stadium downtown. I vaguely understand how text-to-image works, and so it makes sense that the vector space for "blue j…
> sportsball This is not the flex you think it is. You don't have to like sports, but snarking on people who do doesn't make you intellectual, it just makes you come across as a douchebag, no different than a sports fan making fun of "D&D nerds" or something.