Live data from Hacker News

Stable Video Diffusion

stability.ai

101–110 of 316 posts

Re: Stable Video Diffusion

#101
post #69

Earlier quoted context omitted.

> Has anyone come across a solution where model can iterate (eg, with prompts like "move the bicycle to the left side of the photo")? It feels like we're close. I feel like we're close too, but for another reason. For although I love SD and these video examples are great... It's a flawed method: they never get lighting correctly and there are many incoherent things just about everywhere. Any 3D artist or photographer…

Are you working on all that?

Probably not. But there does seem to be a clear path to it.

The main issue is going to be having the right dataset. You basically need to record user actions in something like blender (ie: moving a model of a bike to the left of a scene), match it to a text description of the action (ie; "move bike to the left") and match those to before/after snapshots of the resulting file format.

You need a whole metric fuckton of these.

After that, you train your model to produce those 3d scene files instead of image bitmaps.

You can do this for a lot of other tasks. These general purpose models can learn anything that you can usefully represent in data.

I can imagine AGI being, at least in part, a large set of these purpose trained models. Heck, maybe our brains work this way. When we learn to throw a ball, we train a model in a subset of our brain to do just this and then this model is called on by our general consciousness when needed.

Sorry, I'm just rambling here but its very exciting stuff.

Re: Stable Video Diffusion

#102
post #50

Earlier quoted context omitted.

It's really not. Don't get me wrong, this is insanely cool, but it's still a long way from good enough to be truly disruptive.

In a few years' time, teenagers will be consuming shows and films made by their peers, not by streaming providers. They'll forgive and perhaps even appreciate the technical imperfections for the sake of uncensored, original content that fits perfectly with their cultural identity. Actually, when processing power catches up, I'm expecting a movie engine with well-defined characters, scenes, entities, etc., so people w…

Similar to how all the kids today only play itch.io games thanks to Unity and Unreal dramatically lowering the bar of entry into game development.

Oh wait... No.

All it has done is create an environment where indy games are now assumed to be trash unless proven otherwise, making getting traction as a small developer orders of magnitude harder than it has ever been because their efforts are drowning in a sea of mediocrity.

That same thing is already starting to happen on youtube with AI content, and there's no reason for me to expect this going any other way.

Re: Stable Video Diffusion

#103
post #89
post #76

Earlier quoted context omitted.

No offense, but this is absolutely delusional. As long as people can "clock" content generated from these models, it will be treated by consumers as low-effort drivel, no matter how much actual artistic effort goes in the exercise. Only once these systems push through the threshold of being indistinguishable from artistry will all hell break loose, and we are still very far from that. Paint-by-numbers low-effort mark…

Very far, yes, but also in a fast moving field. CGI in films used to be obvious all the time no matter how good the artists using it, now it's everywhere and only noticeable when that's the point; the gap from Tron to Fellowship of the Ring was 19.5 years. My guess is the analogy here puts the quality of existing genAI somewhere near the equivalent of early TV CGI, given its use in one of the Marvel title sequences e…

something unrelated improved overtime so something else unrelated will also improve to whatever goal you've set in your mind

weird logic circles yall keep making to justify your beliefs, i mean the world is very easy like you just described if you completely strip all nuance and complexity

people used to believe at the start of the space race we'd have mars colonies by now because they looked at the rate of technological advancement from 1910 to 1970, from the first flight to landing on the moon; yet that didn't happen because everything doesn't follow the same repeatable patterns

Re: Stable Video Diffusion

#104

In the video towards the bottom of the page, there are two birds (blue jays), but in the background there are two identical buildings (which look a lot like the CN Tower). CN Tower is the main landmark of Toronto, whose baseball team happens to be the Blue Jays. It's located near the main sportsball stadium downtown. I vaguely understand how text-to-image works, and so it makes sense that the vector space for "blue j…

> Has anyone come across a solution where model can iterate (eg, with prompts like "move the bicycle to the left side of the photo")? It feels like we're close. I feel like we're close too, but for another reason. For although I love SD and these video examples are great... It's a flawed method: they never get lighting correctly and there are many incoherent things just about everywhere. Any 3D artist or photographer…

> you'll describe something and you'll get a full 3D scene, with 3D models, source of lights set up, etc.

I'm always confused why I don't hear more about projects going in this direction. Controlnets are great, but there's still quite a lot of hallucination and other tiny mistakes that a skilled human would never make.

Re: Stable Video Diffusion

#105

Earlier quoted context omitted.

> Has anyone come across a solution where model can iterate (eg, with prompts like "move the bicycle to the left side of the photo")? It feels like we're close. I feel like we're close too, but for another reason. For although I love SD and these video examples are great... It's a flawed method: they never get lighting correctly and there are many incoherent things just about everywhere. Any 3D artist or photographer…

Whats your reasoning for feeling that we're close?

We do it for text, audio and bitmapped images. A 3D scene file format is no different, you could train a model to output a blender file format instead of a bitmap.

It can learn anything you have data for.

Heck, we do it with geospatial data already, generating segmentation vectors. Why not 3D?

Re: Stable Video Diffusion

#106

Looks like I'm still good for my bet with some friends that before 2028 a team of 5-10 people will create a blockbuster style movie that today costs 100+ million USD on a shoestring budget and we won't be able to tell.

I'm imagining more of an AI that takes a standard movie screenplay and a sidecar file, similar to a CSS file for the web and generates the movie. This sidecar file would contain the "director" of the movie, with camera angles, shot length and speed, color grading, etc. Don't like how the new Dune movie looks? Edit the stylesheet and make it your own. Personalized remixed blockbusters.

On a more serious note, I don't think Roger Deakins has anything to worry about right now. Or maybe ever. We've been here before. DAWs opened up an entire world of audio production to people that could afford a laptop and some basic gear. But we certainly do not have a thousand Beatles out there. It still requires talent and effort.

Re: Stable Video Diffusion

#107
post #7

The rate of progress in ML this past year has been breath taking. I can’t wait to see what people do with this once controlnet is properly adapted to video. Generating videos from scratch is cool, but the real utility of this will be the temporal consistency. Getting stable video out of stable diffusion typically involves lots of manual post processing to remove flicker.

What was the big “unlock” that allowed so much progress this past year?

I ask as a noob in this area.

Re: Stable Video Diffusion

#108

Earlier quoted context omitted.

In a few years' time, teenagers will be consuming shows and films made by their peers, not by streaming providers. They'll forgive and perhaps even appreciate the technical imperfections for the sake of uncensored, original content that fits perfectly with their cultural identity. Actually, when processing power catches up, I'm expecting a movie engine with well-defined characters, scenes, entities, etc., so people w…

Similar to how all the kids today only play itch.io games thanks to Unity and Unreal dramatically lowering the bar of entry into game development. Oh wait... No. All it has done is create an environment where indy games are now assumed to be trash unless proven otherwise, making getting traction as a small developer orders of magnitude harder than it has ever been because their efforts are drowning in a sea of medioc…

It took ~2 years for my 10 year old daughter to get bored and give up the shitty user made roblox games and start playing on switch, steam or ps4.

Re: Stable Video Diffusion

#109
post #7

The rate of progress in ML this past year has been breath taking. I can’t wait to see what people do with this once controlnet is properly adapted to video. Generating videos from scratch is cool, but the real utility of this will be the temporal consistency. Getting stable video out of stable diffusion typically involves lots of manual post processing to remove flicker.

What was the big “unlock” that allowed so much progress this past year? I ask as a noob in this area.

Stable diffusion open source release and llama release

Re: Stable Video Diffusion

#110

I admit I'm ignorant about these model's inner workings, but I don't understand why text is the chosen input format for these models. It was the same for image generation, where one needed to produce text prompts to create the image, and stuff like img2img and Controlnet that allowed things like controlling poses and inpainting, or having multiple prompts with masks controlling which part of the image is influenced b…

Imago Deo? The Word is what is spoken when we create.

The input eventually becomes meanings mapped to reality.

Post reply on HN