Live data from Hacker News

Imagen Video: high definition video generation with diffusion models

imagen.research.google

131–140 of 500 posts

Re: Imagen Video: high definition video generation with diffusion models

#131
post #29

Earlier quoted context omitted.

When you animate a horse, does it have 5 legs with weird backwards joints? If not, your job is probably safe for now.

How long do you think until the horse looks perfect? 12 months? 5 years? I’m still 30 and I don’t see how my industry won’t be entirely disrupted by this within the next decade. And that’s my optimistic projection. It could be we have amazing output in 24 months.

IT has been disrupting itself for six decades and there are more developers than ever, with high pay.

Re: Imagen Video: high definition video generation with diffusion models

#132
post #102

These are baby steps towards what I think will be the eventual "disruption" to the film and tv industry. Directors will simply be able to write a script/prompt long enough and detailed enough for something like Imagen (or it's successors) to convert into a feature-length show. Certainly we're very, very far away from that level of cinematic detail and crispness. But I believe that is where this leads... complete with…

I feel like this is very similar to those people who say "have you seen GPT-3? Soon there will be no programmers anymore and all of the code will be generated," and it's wrong for the same reasons.

Can GPT-3 generate good code from vague prompts? Yes, it's surprisingly, sometimes shockingly good at it. Is it ever going to be a replacement for programmers? No, probably not. Same here. This tool's great grandchild is never going to take a rough idea for a movie and churn out a blockbuster film. It'll certainly be a powerful tool in the toolbox of creators, especially the ones on a budget, but it won't make art generation obsolete.

Re: Imagen Video: high definition video generation with diffusion models

#133

It's interesting that these models can generate seemingly anything, but the prompt is taken only as a vague suggestion. From the first 15 examples shown to me, only one contained all elements of the prompt, and it was one of the simplest ("an astronaut riding a horse", versus e.g. "a glass ball falling in water" where it's clear it was a water droplet falling and not a glass ball). We're seeing leaps in random capabi…

In my experience with stable diffusion tools, there is some parameter that specifies how closely you would like it to follow the prompt, which is balanced with giving the AI more freedom to be creative and make the output look better.

Re: Imagen Video: high definition video generation with diffusion models

#134
post #84

Earlier quoted context omitted.

There's a huge gap between "that's pretty cool" and a feature length film. People want to create specific stories with specific scenes in specific places that look a specific way. A "Couple kissing in the rain " prompt isn't going to produce something people are going to pay to see. It's more likely that you're still going to be filming/editing/animating but will have an AI layer on top that produces extra effects or…

I actually think it's the opposite, AI will probably be writing the stories and humans might occasionally film a few scenes. ~95% of TV shows and movies are cookie-cutter content, with cookie-cutter acting and production values, with the same hooks and the same tropes regurgitated over and over again. Heck they can't even figure out how to make new IP so they keep making reruns of the same old stuff like Star Wars, M…

[deleted]

Re: Imagen Video: high definition video generation with diffusion models

#135

Google continues to blow my mind with these models, but I think their ethics strategy is totally misguided and will result in them failing to capture this market. The original Google Search gave similarly never-before-seen capabilities to people, and you could use it for good or bad - Google did not seem to have any ethical concerns around, for example, letting children use their product and come across NSFW content…

Personally, I find it infuriating that Google seems to believe they are the arbiters of morality and truth simply because some of their predecessors figured out good internet search and how to profitably place ads. Google has no special claim to be able to responsibly use these models just because they are rich.

Re: Imagen Video: high definition video generation with diffusion models

#136

"We have decided not to release the Imagen Video model or its source code until these concerns are mitigated" Okay then why even post it in the first place? What exactly is Google going to do with this model?

This whole holier-than-thou moralizing strikes me as trying to steer the conversation away from the real issue, which came into spotlight with Stable Diffusion - one of authorship/violating the IP rights of artists, who now have come down in force against their would be tech overlords who are in the process or repackaging and reselling their work. This forced ideological posturing of 'if we give it to the plebes, the…

> repackaging and reselling their work.

It's not their work unless it's identical, but in practice generated images are substantially different. Drawing in the style of is not copying, it's creative and it also depends on the "dialogue" with the prompter to get to the right image. The artist names added to the prompts act more like landmarks in the latent space, they are a useful shortcut to specifying the style.

If you look at the data itself it's ridiculous - the dataset is 2.3 billion images and the model 4.6 GB, that means it keeps a 2 byte summary from each work it "copies".

Re: Imagen Video: high definition video generation with diffusion models

#137
post #34

I feel like in a not so far future, all this will be generalized into "generate new from all the existing". And at some point later, "all the existing" will be corrupted by the integrated "new" at it will all be chaos. I'm joking, it will be fun all along. :)

It's true, how will future AI train when the training datasets are themselves filled with AI media?

Feedback from whoever is consuming the content it produces.

Re: Imagen Video: high definition video generation with diffusion models

#138

Google continues to blow my mind with these models, but I think their ethics strategy is totally misguided and will result in them failing to capture this market. The original Google Search gave similarly never-before-seen capabilities to people, and you could use it for good or bad - Google did not seem to have any ethical concerns around, for example, letting children use their product and come across NSFW content…

I’ve heard a lot of “data is the new oil” talk and the inevitability of google’s dominance yet I’m inclined to agree with you. Stable diffusion was a big wakeup call where it was clear how much value freedom and creativity really had.

The ethics problem is an artifact of googles model of trying to keep their AI under lock and key and carefully controlled and opaque to outsiders in how the sausage gets made and what it’s made out of. Ultimately I think many of these products will fail because there is a misalignment between what Google thinks you should be able to do with their AI and what people want to do with AI.

Whenever I see an AI ethicists speak I can’t help but think of priests attempting to control the printing press to prevent the spread of dangerous ideas completely sure of their own morality. History will remember them as villains.

Re: Imagen Video: high definition video generation with diffusion models

#139
post #44

Earlier quoted context omitted.

Isn't Imagen a diffusion model? From the abstract: > We present Imagen Video, a text-conditional video generation system based on a cascade of video diffusion models

"Stable Diffusion" is a particular brand from the company Stability AI that is famously open sourcing all of their models.

Pedantically, Stable Diffusion v1.4 is the one model where weights were open sourced and released. Stable Diffusion v1.5, announced September 8th and live on their API, was to be released in "a week or two" but still has yet to be released to the general public.

https://discord.com/channels/1002292111942635562/10022921127...

Re: Imagen Video: high definition video generation with diffusion models

#140

Google continues to blow my mind with these models, but I think their ethics strategy is totally misguided and will result in them failing to capture this market. The original Google Search gave similarly never-before-seen capabilities to people, and you could use it for good or bad - Google did not seem to have any ethical concerns around, for example, letting children use their product and come across NSFW content…

I imagine their lawyers guide them on some of this.
Post reply on HN