Live data from Hacker News

Video to video with Stable Diffusion

stable-diffusion-art.com

31–40 of 110 posts

Re: Video to video with Stable Diffusion

#31
post #7

Not going to bother with this until the temporal cohesion issue is solved. The results are cool, but the variety in frames makes it look like a very specific and distracting art style instead of true animation.

Did you scroll down and see the different methods used, specifically method 5? Not perfect, but getting pretty close.

The light reflection on the shoulder doesn't make much sense. I think this is less pronounced when soft light is involved. I wonder if sharp light and reflections will take longer to be handled properly. Every frame on the shoulder looks ok on its own, but when animated it feels like it was taken in a different environment than the rest.

I suspect depth maps (SD2 already supports them) could be used to achieve that in the future.

I wonder if a diffusion model could accept an "onion skin" noise, so the transitions between frames would be less jarring. Can someone with more knowledge than me explain what's the most promising approach here?

Re: Video to video with Stable Diffusion

#32
post #23
post #19

Earlier quoted context omitted.

The poster is just saying that "the "glitch" of today will be the nostalgic style of tomorrow".

Exactly. I have no idea what the other commenters thought I meant.

They are saying that there is no need to figure out how to recreate "the "glitch" of today that will be the nostalgic style of tomorrow" as the "old"/current models are all available on GitHub and huggingface, and will be in the future.

There is no need to figure out how to recreate this style as you will be able to find it on these platforms in the future. The commenter understood you, but you did not understand the commenter.

Re: Video to video with Stable Diffusion

#33
post #30
post #7

Not going to bother with this until the temporal cohesion issue is solved. The results are cool, but the variety in frames makes it look like a very specific and distracting art style instead of true animation.

> the variety in frames makes it look like a very specific and distracting art style instead of true animation. Well, what is "true animation"? There are number of types of traditional animation where you draw both the characters and the backgrounds from scratch for every frame, so you get this variety naturally. Or you have more practical techniques like strata-cut animation[1], where again you get what we might thi…

In this case we have objects disappearing and reappearing, and the design of clothing and objects (headphones/tiara) being significantly different from frame to frame, pieces of sleeves blinking in and out of existence, headphones becoming tiaras becoming helmets... I'm sure you can find examples of such things in traditional animation but there's clearly something peculiar going on here due to the frames being processed mostly standalone, and it's clearly specific to this exact way of processing. Even if you "draw both the characters and the backgrounds from scratch for every frame", normally the animator would at least have the previous frames and/or overall design at his disposal as a reference. What's happening here is more like those crowdsourcing experiments where you let many people draw a single frame without knowing what the others are up to.

Re: Video to video with Stable Diffusion

#35

Earlier quoted context omitted.

Did you scroll down and see the different methods used, specifically method 5? Not perfect, but getting pretty close.

The light reflection on the shoulder doesn't make much sense. I think this is less pronounced when soft light is involved. I wonder if sharp light and reflections will take longer to be handled properly. Every frame on the shoulder looks ok on its own, but when animated it feels like it was taken in a different environment than the rest. I suspect depth maps (SD2 already supports them) could be used to achieve that i…

>The light reflection on the shoulder doesn't make much sense.

People will really grasp at anything to find something to critique in generative AI, huh?

Re: Video to video with Stable Diffusion

#36
I was going to get a 4090 this year, but I just don't think 24gb VRAM is going to bee enough in the short term future(for AI related stuff).

Ended up getting a 3060 for 1/4 of the cost, and I'm planning on using/paying for Colab until some 6090 comes out with 128gb vram.

Something that kind of bothers me that I don't understand. Why is there such an obsession about having small computers/servers? I don't care if I have to put my computer in the basement because its the size of a closet or two. Maybe there is an EE issue, like too much current or EMF if you make a large computer.

I'd usually use colab, for 99% of the jobs... but I admittedly want to do some nsfw stuff with me and the wife with AI Art.

Re: Video to video with Stable Diffusion

#37
post #30

Earlier quoted context omitted.

> the variety in frames makes it look like a very specific and distracting art style instead of true animation. Well, what is "true animation"? There are number of types of traditional animation where you draw both the characters and the backgrounds from scratch for every frame, so you get this variety naturally. Or you have more practical techniques like strata-cut animation[1], where again you get what we might thi…

In this case we have objects disappearing and reappearing, and the design of clothing and objects (headphones/tiara) being significantly different from frame to frame, pieces of sleeves blinking in and out of existence, headphones becoming tiaras becoming helmets... I'm sure you can find examples of such things in traditional animation but there's clearly something peculiar going on here due to the frames being proce…

you'd think a real AI would be aware of all these issues and not let this happen.

Re: Video to video with Stable Diffusion

#38
post #23

Earlier quoted context omitted.

Exactly. I have no idea what the other commenters thought I meant.

They are saying that there is no need to figure out how to recreate "the "glitch" of today that will be the nostalgic style of tomorrow" as the "old"/current models are all available on GitHub and huggingface, and will be in the future. There is no need to figure out how to recreate this style as you will be able to find it on these platforms in the future. The commenter understood you, but you did not understand the…

Yes. Because of the way reply notifications come in I thought it was a reply to a different comment.

But I still think that's missing the point of my (not entirely serious) comment.

People still have access to analogue amplifiers but digital simulations of them are still developed.

As models and workflows improve, if people want a flickery old-school look then they may well simulate it rather than go through the hassle of running old tools that might not mesh well with newer workflows.

Re: Video to video with Stable Diffusion

#39

I was going to get a 4090 this year, but I just don't think 24gb VRAM is going to bee enough in the short term future(for AI related stuff). Ended up getting a 3060 for 1/4 of the cost, and I'm planning on using/paying for Colab until some 6090 comes out with 128gb vram. Something that kind of bothers me that I don't understand. Why is there such an obsession about having small computers/servers? I don't care if I ha…

> Why is there such an obsession about having small computers/servers

More likely just cost. If a "small" computer costs $2k, how much do you think a "big" one will cost? $10k+ probably, what is the market for that?

And you can buy a rack in your home and put 10 servers in it if you really want a closet sized computer.

Re: Video to video with Stable Diffusion

#40
post #7

Not going to bother with this until the temporal cohesion issue is solved. The results are cool, but the variety in frames makes it look like a very specific and distracting art style instead of true animation.

I've had the idea for a long time but never managed to try it: Alias-free convolutions (like StyleGAN3) may help with temporal cohesion.

The StyleGAN3 project page shows some good videos: https://nvlabs.github.io/stylegan3/

Post reply on HN