Live data from Hacker News

Imagen Video: high definition video generation with diffusion models

imagen.research.google

181–190 of 500 posts

Re: Imagen Video: high definition video generation with diffusion models

#181

It's interesting that these models can generate seemingly anything, but the prompt is taken only as a vague suggestion. From the first 15 examples shown to me, only one contained all elements of the prompt, and it was one of the simplest ("an astronaut riding a horse", versus e.g. "a glass ball falling in water" where it's clear it was a water droplet falling and not a glass ball). We're seeing leaps in random capabi…

In my experience with stable diffusion tools, there is some parameter that specifies how closely you would like it to follow the prompt, which is balanced with giving the AI more freedom to be creative and make the output look better.

Yes, that might be the case. Though the prompts don't seem to try showcasing model creativity, so I'd be surprised if Google picked a temperature so high that it significantly deviated from the prompt so often.

Re: Imagen Video: high definition video generation with diffusion models

#182

Google continues to blow my mind with these models, but I think their ethics strategy is totally misguided and will result in them failing to capture this market. The original Google Search gave similarly never-before-seen capabilities to people, and you could use it for good or bad - Google did not seem to have any ethical concerns around, for example, letting children use their product and come across NSFW content…

I will say, I've enjoyed playing with stable diffusion, I've been impressed with the explosion of tools built around it, and the stuff people are creating ... But all the stuff about bias in data is true. It really likes to render white people, unless you really specifically tell it something else ... in which case, you may receive an exaggerated stereotype. It seems to like producing younger adults. If all stock photography tomorrow forward was replaced with stable diffusion images, even ignoring the weird bodies and messed up faces and stuff, I think it would create negative effects. And once models are naively trained on images produced by the previous generation, how much worse will it be?

I don't think "don't let the plebes have the models" is a good stance. But neither is pretending that the ethics and bias issues aren't here.

Re: Imagen Video: high definition video generation with diffusion models

#183
post #171

Earlier quoted context omitted.

I feel like this is very similar to those people who say "have you seen GPT-3? Soon there will be no programmers anymore and all of the code will be generated," and it's wrong for the same reasons. Can GPT-3 generate good code from vague prompts? Yes, it's surprisingly, sometimes shockingly good at it. Is it ever going to be a replacement for programmers? No, probably not. Same here. This tool's great grandchild is n…

> This tool's great grandchild is never going to take a rough idea for a movie and churn out a blockbuster film. What about the tool's nth child though? I think saying it will never do it is a bit much, given what we know about human ingenuity and economic incentives.

I think individual special effects sound very plausible. "Okay, robot, make it so that his arm gets vaporized by an incoming laser, kinda like the same effect in Iron Man 7" is believable to me.

But ultimately these things copy other stuff. Artists are often trying to create something that is, at least a bit, new. New is where this approach falls over. By its nature, these things paint from examples. They can design Rococo things because they have seen many Rococo things and know what the word means. But they can't come up with a new style and use it consistently. "Make a video game with a fun and unique mechanic" is not something these things could ever do.

I think it's certainly possible, maybe inevitable, that some AI system in the distant future could do that, but it won't be based on this style of algorithm. An algorithm that can take "make a fun romantic comedy with themes of loneliness" and make something award worthy will be a lot closer to AGI than it will be to this stuff.

Re: Imagen Video: high definition video generation with diffusion models

#184

Earlier quoted context omitted.

Why would I want to watch AI-generated content?

Procedurally generated games can be quite fun, if AI content gets good enough, why wouldn't you want to watch it?

Because anything that an AI can produce, no matter how "intrinsically" good, becomes trivial, tedious and with zero value (both economic and general).

Re: Imagen Video: high definition video generation with diffusion models

#185
post #115

I’ll be honest, as someone who worked in the film industry for a decade, this thread is depressing. It’s not the technology, it’s all the people in these comments who have never worked in the industry clamouring for its demise. One could brush it off as tech heads being over exuberant, but it’s the lack of understanding of how much fine control goes into each and every shot of a film that is depressing. If I, as a cr…

I agree with you, but I wouldn't take it so personally. There have been people claiming machines will make one industry or another obsolete for as long as we've had machines. In a way, sometimes they're right! But this doesn't mean the people are obsolete. Excel never made accountants obsolete, it just made their jobs easier and less tedious. I feel like content generation tools might offer something similar. How nic…

Oh I don’t take it personally so much as I find it sad how quickly people in the tech sphere are so quick to extol the virtues of things they have no familiarity with.

Every AI art thread is full of people who have clearly never attempted to make professional art commenting as if they’re experts in the domain

Re: Imagen Video: high definition video generation with diffusion models

#186
post #115

I’ll be honest, as someone who worked in the film industry for a decade, this thread is depressing. It’s not the technology, it’s all the people in these comments who have never worked in the industry clamouring for its demise. One could brush it off as tech heads being over exuberant, but it’s the lack of understanding of how much fine control goes into each and every shot of a film that is depressing. If I, as a cr…

The term "creative" is so pretentious, as if only content generation involves creativity. Your post reminds me of all the photographers that said digital photography would remain niche and never replace film. The current models are toys made by small groups. It's not hard to imagine AI generated film being much more compelling when the entire industry of engineers and "creatives" refine and evolve the ecosystem to ta…

Why is it any more pretentious than “developer” or “engineer”?

Also businesses don’t always go for cheaper. They go for maximum ROI.

I’ve worked on tons of marvel films for example, and I quite well know where AI fits and speeds things up. I also know where client studios will pay a pretty penny for more art directed results rather than going for the cheapest vendor.

Re: Imagen Video: high definition video generation with diffusion models

#187

Earlier quoted context omitted.

Emad (founder of Stability AI) has said they already have video model training underway, as well as text and audio. Exciting times.

Is this going to end up into a single model, where its trained on text and images and audio and videos and 3d models, and it can do anything to anything depending on what you ask of it? Feels like the cross-training would help yield stronger results.

Probably not. We're actually headed towards many smaller models that call each other, because VRAM is the limiting factor in application, and if the domains aren't totally dependent on each other it's easier to have one model produce bad output, then detect that bad output and feed it into another model that cleans up the problem (like fixing faces in stable diffusion output).

The human brain is modularized like this, so I don't think it'll be a limitation.

Re: Imagen Video: high definition video generation with diffusion models

#188

Google continues to blow my mind with these models, but I think their ethics strategy is totally misguided and will result in them failing to capture this market. The original Google Search gave similarly never-before-seen capabilities to people, and you could use it for good or bad - Google did not seem to have any ethical concerns around, for example, letting children use their product and come across NSFW content…

>You can't type any prompt that's "unsafe", you can't generate images of people, there are so many stupid limitations that the product is practically useless other than niche scenarios

Imagen and Imagen Video is not released to the public at all. You might be confusing it with OpenAI's models.

Re: Imagen Video: high definition video generation with diffusion models

#189

Google continues to blow my mind with these models, but I think their ethics strategy is totally misguided and will result in them failing to capture this market. The original Google Search gave similarly never-before-seen capabilities to people, and you could use it for good or bad - Google did not seem to have any ethical concerns around, for example, letting children use their product and come across NSFW content…

Google has no practical way to address ethics at Google-scale. Their ability to operate at all depends as ever upon outsourcing ethics to machine learning algorithms.

Re: Imagen Video: high definition video generation with diffusion models

#190

I really like these videos because they're trippy. Someone should work on a neural net to generate trippy videos. It would probably be much easier than realistic videos (esp. because these videos are noticeably generated from obvious to subtle). Also is nobody paying attention to the fact that they got words correct? At least "Imagen Video". Prior models all suck at word order

Both models, imagen and parti didn't had a problem with text. Only dalle and stable diffusion
Post reply on HN