I'm still puzzled as to how these "non-commercial" model licenses are supposed to be enforceable. Software licenses govern the redistribution of the software , not products produced with it. An image isn't GPL'd because it was produced with GIMP.
Visual Studio Community (and many other products) only allows "non-commercial" usage. Sounds like it limits what you can do with what you produce with it. At the end of the day, a license is a legal contract. If you agree that an image which you produce with some software will be GPL'ed, it's enforceable. As an example, see the Creative Commons license, ShareAlike clause: > If you remix, transform, or build upon the…
Stable Video Diffusion
171–180 of 316 posts
Re: Stable Video Diffusion
#172In the video towards the bottom of the page, there are two birds (blue jays), but in the background there are two identical buildings (which look a lot like the CN Tower). CN Tower is the main landmark of Toronto, whose baseball team happens to be the Blue Jays. It's located near the main sportsball stadium downtown. I vaguely understand how text-to-image works, and so it makes sense that the vector space for "blue j…
> Has anyone come across a solution where model can iterate (eg, with prompts like "move the bicycle to the left side of the photo")? It feels like we're close. I feel like we're close too, but for another reason. For although I love SD and these video examples are great... It's a flawed method: they never get lighting correctly and there are many incoherent things just about everywhere. Any 3D artist or photographer…
Re: Stable Video Diffusion
#173Earlier quoted context omitted.
Whats your reasoning for feeling that we're close?
We do it for text, audio and bitmapped images. A 3D scene file format is no different, you could train a model to output a blender file format instead of a bitmap. It can learn anything you have data for. Heck, we do it with geospatial data already, generating segmentation vectors. Why not 3D?
A 3D scene is vastly more complex, and the way you consume it is tangential to the rendering of it we use to interpret. It is a collection of arbitrary data structures.
We’ll need a new approach for this kind of problem
Re: Stable Video Diffusion
#174Re: Stable Video Diffusion
#175Earlier quoted context omitted.
Whats your reasoning for feeling that we're close?
We do it for text, audio and bitmapped images. A 3D scene file format is no different, you could train a model to output a blender file format instead of a bitmap. It can learn anything you have data for. Heck, we do it with geospatial data already, generating segmentation vectors. Why not 3D?
Re: Stable Video Diffusion
#176Earlier quoted context omitted.
Visual Studio Community (and many other products) only allows "non-commercial" usage. Sounds like it limits what you can do with what you produce with it. At the end of the day, a license is a legal contract. If you agree that an image which you produce with some software will be GPL'ed, it's enforceable. As an example, see the Creative Commons license, ShareAlike clause: > If you remix, transform, or build upon the…
Do you have link for the VS Community terms you're describing? What I've found is directly contradictory: "Any individual developer can use Visual Studio Community to create their own free or paid apps." From https://visualstudio.microsoft.com/vs/community/
> In enterprise organizations (meaning those with >250 PCs or >$1 Million US Dollars in annual revenue), no use is permitted beyond the open source, academic research, and classroom learning environment scenarios described above.
Re: Stable Video Diffusion
#177Earlier quoted context omitted.
We do it for text, audio and bitmapped images. A 3D scene file format is no different, you could train a model to output a blender file format instead of a bitmap. It can learn anything you have data for. Heck, we do it with geospatial data already, generating segmentation vectors. Why not 3D?
Text, audio, and bitmapped images are data. Numbers and tokens. A 3D scene is vastly more complex, and the way you consume it is tangential to the rendering of it we use to interpret. It is a collection of arbitrary data structures. We’ll need a new approach for this kind of problem
> A 3D scene is vastly more complex
3D scenes, in fact, are also data, numbers and tokens. (Well, numbers, but so are tokens.)
Re: Stable Video Diffusion
#178Earlier quoted context omitted.
> Has anyone come across a solution where model can iterate (eg, with prompts like "move the bicycle to the left side of the photo")? It feels like we're close. I feel like we're close too, but for another reason. For although I love SD and these video examples are great... It's a flawed method: they never get lighting correctly and there are many incoherent things just about everywhere. Any 3D artist or photographer…
> you'll describe something and you'll get a full 3D scene, with 3D models, source of lights set up, etc. I'm always confused why I don't hear more about projects going in this direction. Controlnets are great, but there's still quite a lot of hallucination and other tiny mistakes that a skilled human would never make.
Probably because they aren't as advanced and the demos aren't as impressive to nontechnical audiences who don't understand the implications: there’s lots of work on text-to-3d-model generation, and even plugins for some stable diffusion UIs (e.g., MotionDiff for ComyUI.)
Re: Stable Video Diffusion
#179Earlier quoted context omitted.
Controlnet is adapted to video today, the issues are that it's very slow. Haven't you seen the insane quality of videos on civitai?
> Haven't you seen the insane quality of videos on civitai? I have not, so I went to https://civitai.com/ which I guess is what you're talking about? But I cannot find a single video there, just images and models.
https://www.youtube.com/watch?v=3WWy98ylLT4
https://www.youtube.com/shorts/1vqOjYWEF84
https://www.youtube.com/shorts/jOIb9QbrhZ8
https://www.youtube.com/shorts/C3F_YI84TXA
https://www.youtube.com/shorts/4IqJHozY4F0
https://www.youtube.com/shorts/h3OmBLlm5-g
https://www.youtube.com/shorts/ZT7tuIgSDRk
https://www.youtube.com/shorts/WnUYbsOMyvs
https://www.youtube.com/shorts/BKKqX2aMlSg
The inconsistencies are what's most interesting in these videos in fact
Re: Stable Video Diffusion
#180For example, the man in the cowboy hat seems he is almost gagging. In the train video the tracks seem to be too wide while the train ice skates across them.