Live data from Hacker News

Stable Video Diffusion

stability.ai

171–180 of 316 posts

Re: Stable Video Diffusion

#171

I'm still puzzled as to how these "non-commercial" model licenses are supposed to be enforceable. Software licenses govern the redistribution of the software , not products produced with it. An image isn't GPL'd because it was produced with GIMP.

Visual Studio Community (and many other products) only allows "non-commercial" usage. Sounds like it limits what you can do with what you produce with it. At the end of the day, a license is a legal contract. If you agree that an image which you produce with some software will be GPL'ed, it's enforceable. As an example, see the Creative Commons license, ShareAlike clause: > If you remix, transform, or build upon the…

Do you have link for the VS Community terms you're describing? What I've found is directly contradictory: "Any individual developer can use Visual Studio Community to create their own free or paid apps." From https://visualstudio.microsoft.com/vs/community/

Re: Stable Video Diffusion

#172

In the video towards the bottom of the page, there are two birds (blue jays), but in the background there are two identical buildings (which look a lot like the CN Tower). CN Tower is the main landmark of Toronto, whose baseball team happens to be the Blue Jays. It's located near the main sportsball stadium downtown. I vaguely understand how text-to-image works, and so it makes sense that the vector space for "blue j…

> Has anyone come across a solution where model can iterate (eg, with prompts like "move the bicycle to the left side of the photo")? It feels like we're close. I feel like we're close too, but for another reason. For although I love SD and these video examples are great... It's a flawed method: they never get lighting correctly and there are many incoherent things just about everywhere. Any 3D artist or photographer…

This isn't coming, it's already here. https://github.com/gsgen3d/gsgen Yes, it's just 3D models for now, but it can do whole scenes generations, it's just not great yet at it. The tech is there but just need to improve.

Re: Stable Video Diffusion

#173
post #105

Earlier quoted context omitted.

Whats your reasoning for feeling that we're close?

We do it for text, audio and bitmapped images. A 3D scene file format is no different, you could train a model to output a blender file format instead of a bitmap. It can learn anything you have data for. Heck, we do it with geospatial data already, generating segmentation vectors. Why not 3D?

Text, audio, and bitmapped images are data. Numbers and tokens.

A 3D scene is vastly more complex, and the way you consume it is tangential to the rendering of it we use to interpret. It is a collection of arbitrary data structures.

We’ll need a new approach for this kind of problem

Re: Stable Video Diffusion

#175
post #105

Earlier quoted context omitted.

Whats your reasoning for feeling that we're close?

We do it for text, audio and bitmapped images. A 3D scene file format is no different, you could train a model to output a blender file format instead of a bitmap. It can learn anything you have data for. Heck, we do it with geospatial data already, generating segmentation vectors. Why not 3D?

We do it for 3D, too.

https://guytevet.github.io/mdm-page/

Re: Stable Video Diffusion

#176

Earlier quoted context omitted.

Visual Studio Community (and many other products) only allows "non-commercial" usage. Sounds like it limits what you can do with what you produce with it. At the end of the day, a license is a legal contract. If you agree that an image which you produce with some software will be GPL'ed, it's enforceable. As an example, see the Creative Commons license, ShareAlike clause: > If you remix, transform, or build upon the…

Do you have link for the VS Community terms you're describing? What I've found is directly contradictory: "Any individual developer can use Visual Studio Community to create their own free or paid apps." From https://visualstudio.microsoft.com/vs/community/

Enterprise organizations are not allowed to use VS Community for commercial purposes:

> In enterprise organizations (meaning those with >250 PCs or >$1 Million US Dollars in annual revenue), no use is permitted beyond the open source, academic research, and classroom learning environment scenarios described above.

Re: Stable Video Diffusion

#177
post #105

Earlier quoted context omitted.

We do it for text, audio and bitmapped images. A 3D scene file format is no different, you could train a model to output a blender file format instead of a bitmap. It can learn anything you have data for. Heck, we do it with geospatial data already, generating segmentation vectors. Why not 3D?

Text, audio, and bitmapped images are data. Numbers and tokens. A 3D scene is vastly more complex, and the way you consume it is tangential to the rendering of it we use to interpret. It is a collection of arbitrary data structures. We’ll need a new approach for this kind of problem

> Text, audio, and bitmapped images are data. Numbers and tokens.

> A 3D scene is vastly more complex

3D scenes, in fact, are also data, numbers and tokens. (Well, numbers, but so are tokens.)

Re: Stable Video Diffusion

#178
post #104

Earlier quoted context omitted.

> Has anyone come across a solution where model can iterate (eg, with prompts like "move the bicycle to the left side of the photo")? It feels like we're close. I feel like we're close too, but for another reason. For although I love SD and these video examples are great... It's a flawed method: they never get lighting correctly and there are many incoherent things just about everywhere. Any 3D artist or photographer…

> you'll describe something and you'll get a full 3D scene, with 3D models, source of lights set up, etc. I'm always confused why I don't hear more about projects going in this direction. Controlnets are great, but there's still quite a lot of hallucination and other tiny mistakes that a skilled human would never make.

> I'm always confused why I don't hear more about projects going in this direction.

Probably because they aren't as advanced and the demos aren't as impressive to nontechnical audiences who don't understand the implications: there’s lots of work on text-to-3d-model generation, and even plugins for some stable diffusion UIs (e.g., MotionDiff for ComyUI.)

Re: Stable Video Diffusion

#179

Earlier quoted context omitted.

Controlnet is adapted to video today, the issues are that it's very slow. Haven't you seen the insane quality of videos on civitai?

> Haven't you seen the insane quality of videos on civitai? I have not, so I went to https://civitai.com/ which I guess is what you're talking about? But I cannot find a single video there, just images and models.

https://www.youtube.com/shorts/ZN-NbdFwfNQ

https://www.youtube.com/watch?v=3WWy98ylLT4

https://www.youtube.com/shorts/1vqOjYWEF84

https://www.youtube.com/shorts/jOIb9QbrhZ8

https://www.youtube.com/shorts/C3F_YI84TXA

https://www.youtube.com/shorts/4IqJHozY4F0

https://www.youtube.com/shorts/h3OmBLlm5-g

https://www.youtube.com/shorts/ZT7tuIgSDRk

https://www.youtube.com/shorts/WnUYbsOMyvs

https://www.youtube.com/shorts/BKKqX2aMlSg

The inconsistencies are what's most interesting in these videos in fact

Re: Stable Video Diffusion

#180
Much like in static images, the subtle unintended imperfections are quite interesting to observe.

For example, the man in the cowboy hat seems he is almost gagging. In the train video the tracks seem to be too wide while the train ice skates across them.

Post reply on HN