Very soon, we will be able to change story line of a web series dynamically, a little more thrill, a little more comedy, changing character face to matching ours and others, all in 3D with 360 degree view, how far are we from this ? 5 year ?
Stable Video Diffusion
291–300 of 316 posts
Re: Stable Video Diffusion
#292Earlier quoted context omitted.
There are a few problems: 1) You and I invent our own private "copyright" for data (which is not copyrightable) 2) Everything is fine until my wife walks up to my computer and makes a copy of the data. She's not bound by our private "copyright." She doesn't even know it exists, and shares the data with her bestie. And... our private pseudo-copyright is dead. Also: Licenses are not the same as contracts. There are tim…
> my wife walks up to my computer and makes a copy of the data As you agreed to in our contract, you now need to compensate me for the damage caused by your failure to prevent unauthorized third-party access. Of course you're free to attempt to recover the sum you have to pay me from your wife. > The output of a program is rarely copyrightable by the author (as opposed to the user). The author of the program can make…
1) You have a $1T product.
2) My wife leaks it, or a burglar does. I am a typical consumer, with say, a $20k net worth.
You have two choices:
1) Sue me, recover $20k, and be down $1T (minus $20k, plus litigation fees), and get the press of ruining the life of some innocent random person
2) Not sue me. Be down $1T (including the $20k) .
And yes, the author of a program can put whatever conditions they want into the license: "By using this program, you agree to transfer $1M into my bank account in bit coin, to give me your first-born baby, to swear fealty to me, and to give me your wife it servitude." A court can then read those conditions, have a good laugh, and not enforce them. There are very clear limits on what a court will enforce in licenses (and contracts), and owning the output of a program, and barring exceptional circumstance, courts will not enforce them:
https://www.lexology.com/library/detail.aspx?g=eb52567a-2104...
This is why programmers should learn basic law, not treat is as computer code, and consult lawyers when issues come up. Read by a lawyer, a license or contract with an unenforceable clause is as good as having no such clause.
Re: Stable Video Diffusion
#293In the video towards the bottom of the page, there are two birds (blue jays), but in the background there are two identical buildings (which look a lot like the CN Tower). CN Tower is the main landmark of Toronto, whose baseball team happens to be the Blue Jays. It's located near the main sportsball stadium downtown. I vaguely understand how text-to-image works, and so it makes sense that the vector space for "blue j…
As for your last question yes that exists. There are two models from Meta that do exactly this, instruction based iteration on photos, Emu Edit[0], and videos, Emu Video[1].
There's also LLaVa-interactive[2] for photos where you can even chat with the model about the current image.
[0]: https://emu-edit.metademolab.com/
Re: Stable Video Diffusion
#294Earlier quoted context omitted.
> Has anyone come across a solution where model can iterate (eg, with prompts like "move the bicycle to the left side of the photo")? It feels like we're close. I feel like we're close too, but for another reason. For although I love SD and these video examples are great... It's a flawed method: they never get lighting correctly and there are many incoherent things just about everywhere. Any 3D artist or photographer…
Nah I disagree, this feels like a glorification of the process not the end result. Just because having the 3D model in the scene with all the lighting makes the end result feel more solid to you because you feel you can see the work that's going into it. In the end diffusion technology can make a more realistic image faster than a rendering engine can. I feel pretty strongly that this pipeline will be the foundation…
Re: Stable Video Diffusion
#295Earlier quoted context omitted.
> Haven't you seen the insane quality of videos on civitai? I have not, so I went to https://civitai.com/ which I guess is what you're talking about? But I cannot find a single video there, just images and models.
https://www.youtube.com/shorts/ZN-NbdFwfNQ https://www.youtube.com/watch?v=3WWy98ylLT4 https://www.youtube.com/shorts/1vqOjYWEF84 https://www.youtube.com/shorts/jOIb9QbrhZ8 https://www.youtube.com/shorts/C3F_YI84TXA https://www.youtube.com/shorts/4IqJHozY4F0 https://www.youtube.com/shorts/h3OmBLlm5-g https://www.youtube.com/shorts/ZT7tuIgSDRk https://www.youtube.com/shorts/WnUYbsOMyvs https://www.youtube.com/shorts/B…
Re: Stable Video Diffusion
#296I think this will really open new ways and new doors to creativity and creative expression.
Re: Stable Video Diffusion
#297In the video towards the bottom of the page, there are two birds (blue jays), but in the background there are two identical buildings (which look a lot like the CN Tower). CN Tower is the main landmark of Toronto, whose baseball team happens to be the Blue Jays. It's located near the main sportsball stadium downtown. I vaguely understand how text-to-image works, and so it makes sense that the vector space for "blue j…
Adobe is doing some great work here in my opinion in terms of building AI tools that make sense for artist workflows. This "sneak peak" demo from the recent Adobe Max conference is pretty much exactly what you described, actually better because you can just click on an object in the image and drag it. See video: https://www.adobe.com/max/2023/sessions/project-stardust-gs6...
Re: Stable Video Diffusion
#298Earlier quoted context omitted.
> Has anyone come across a solution where model can iterate (eg, with prompts like "move the bicycle to the left side of the photo")? It feels like we're close. I feel like we're close too, but for another reason. For although I love SD and these video examples are great... It's a flawed method: they never get lighting correctly and there are many incoherent things just about everywhere. Any 3D artist or photographer…
Nah I disagree, this feels like a glorification of the process not the end result. Just because having the 3D model in the scene with all the lighting makes the end result feel more solid to you because you feel you can see the work that's going into it. In the end diffusion technology can make a more realistic image faster than a rendering engine can. I feel pretty strongly that this pipeline will be the foundation…
Re: Stable Video Diffusion
#299Earlier quoted context omitted.
I think the bottleneck is data For single 3D object the biggest dataset is ObjaverseXL with 10M samples For full 3D scenes you could at best get ~1000 scenes with datasets like ScanNet I guess Text2Image models are trained on datasets with 5 billion samples
Oh, I don't know about that. Working in feature film animation, studios have gargantuan model libraries from current and past projects, with a good number (over half) never used by a production but created as part of some production's world building. Plus, generative modeling has been very popular for quite a few years. I don't think getting more 3D models then they could use is a real issue for anyone serious.
Got a public dataset?
Re: Stable Video Diffusion
#300Earlier quoted context omitted.
Nah I disagree, this feels like a glorification of the process not the end result. Just because having the 3D model in the scene with all the lighting makes the end result feel more solid to you because you feel you can see the work that's going into it. In the end diffusion technology can make a more realistic image faster than a rendering engine can. I feel pretty strongly that this pipeline will be the foundation…
I find that very unlikely. LLMs seem capable of simulating human intuition, but not great at simulating real complex physics. Human intuition of how a scene “should” look isn’t always the effect you want to create, and is rarely accurate im guessing
Diffusion models aren't LLMs (they may use something similar as their text encoder layer) and they simulate their training corpus, which usually isn't selected solely for physical fidelity, because that's not actually the single criteria for visual imagery outside of what is created by diffusion models.