Live data from Hacker News

Imagen Video: high definition video generation with diffusion models

imagen.research.google

121–130 of 500 posts

Re: Imagen Video: high definition video generation with diffusion models

#122
post #116
post #103

Earlier quoted context omitted.

I really doubt you’d be able to have the fine grained control that most high end creatives want with any of these diffusion models, let alone the ability to convey specific emotions. At that point, we’d have reached some kind of AI singularity and the disruption would be everywhere not just in the creative sphere

There's no doubt that it's only a matter of time. Like bloggers had the opportunity to compete with newspapers, the ability to generate videos will allow to compete with movies/marvel/netflix/disney & company. Eventually, only high quality content will justify the need to pay for a ticket or a subscription, and there's going to be a lot of free content to watch, with 1000x more people able to publish their ideas, as…

You’re conflating the ability to make things for the masses and being able to automatically generate it.

Film production is already commoditized and anyone can make high end content.

Being able to automatically create that is a different argument than what you posit.

Re: Imagen Video: high definition video generation with diffusion models

#123
post #103
post #102

These are baby steps towards what I think will be the eventual "disruption" to the film and tv industry. Directors will simply be able to write a script/prompt long enough and detailed enough for something like Imagen (or it's successors) to convert into a feature-length show. Certainly we're very, very far away from that level of cinematic detail and crispness. But I believe that is where this leads... complete with…

I really doubt you’d be able to have the fine grained control that most high end creatives want with any of these diffusion models, let alone the ability to convey specific emotions. At that point, we’d have reached some kind of AI singularity and the disruption would be everywhere not just in the creative sphere

I disagree. It's a rudimentary features of all these models to take a concept picture and refine it. It won't be like the director would give a prompt and get a feature length movie, it will be more like the director uses MS Paint (as in a common software for non tech people) to make a scene outline and directs AI to make a stylish and animated version of that. Something is wrong? just erase it and try again. Dalle2 had this interface from the get go. The models just haven't gotten there yet.

Re: Imagen Video: high definition video generation with diffusion models

#124
post #123
post #103

Earlier quoted context omitted.

I really doubt you’d be able to have the fine grained control that most high end creatives want with any of these diffusion models, let alone the ability to convey specific emotions. At that point, we’d have reached some kind of AI singularity and the disruption would be everywhere not just in the creative sphere

I disagree. It's a rudimentary features of all these models to take a concept picture and refine it. It won't be like the director would give a prompt and get a feature length movie, it will be more like the director uses MS Paint (as in a common software for non tech people) to make a scene outline and directs AI to make a stylish and animated version of that. Something is wrong? just erase it and try again. Dalle2…

Try again and do what? How are you directing the shot? How do you erase an emotion? How do you erase and redo inner turmoil when delivering a performance?

Re: Imagen Video: high definition video generation with diffusion models

#125
It's interesting that these models can generate seemingly anything, but the prompt is taken only as a vague suggestion.

From the first 15 examples shown to me, only one contained all elements of the prompt, and it was one of the simplest ("an astronaut riding a horse", versus e.g. "a glass ball falling in water" where it's clear it was a water droplet falling and not a glass ball).

We're seeing leaps in random capabilities (motion! 3D! inpainting! voice editing!), so I wonder if complete prompt accuracy is 3 months or 3 years away. But I wouldn't bet on any longer than that.

Re: Imagen Video: high definition video generation with diffusion models

#126

Google continues to blow my mind with these models, but I think their ethics strategy is totally misguided and will result in them failing to capture this market. The original Google Search gave similarly never-before-seen capabilities to people, and you could use it for good or bad - Google did not seem to have any ethical concerns around, for example, letting children use their product and come across NSFW content…

Another way to look at it is the people at Google are all now quasi-retired with kids and wouldn't be so mad if some scrappy startups ate their business lunches (while they are at home with their fams). Perhaps they are just subsidizing research.

Re: Imagen Video: high definition video generation with diffusion models

#127
The progress of content generation is disorienting! I remember studying Markov Chains and Hidden Markov Models for text generation. Then we had Recurrent Networks which went from LSTMs to Transformers now. At this point we can have a sustained pseudo conversation with a model, which will do trivial tasks for us from a text corpus.

Separately for images we had convolutional networks and Generative Adversarial Networks. Now diffusion models are apparently doing what Transformers did to natural language processing.

In my field, we use shallower feed-forward networks for control using low-dimensional sensor data (for speed & interpretability). Physical constraints (and good-enoughness of classical approaches) make such massive leaps in performance rarer events.

Re: Imagen Video: high definition video generation with diffusion models

#128

Google continues to blow my mind with these models, but I think their ethics strategy is totally misguided and will result in them failing to capture this market. The original Google Search gave similarly never-before-seen capabilities to people, and you could use it for good or bad - Google did not seem to have any ethical concerns around, for example, letting children use their product and come across NSFW content…

“But then the inevitable might occur!” — someone at Google probably.

Re: Imagen Video: high definition video generation with diffusion models

#130

Earlier quoted context omitted.

I feel stupid what are those ethical implications? It seems like just a cool technology to me.

Top two comments are creatives wondering about their future jobs. Ai ethicists have brought up concerns regarding intentional misuse like misinformation. The technology is super cool. Cat is out of the bag. Just like we couldn't really make cryptography illegal, this stuff shouldn't be either. But I dislike how everyone is pretending that AI ethicists and others are completely unfounded just because it is popular to…

> Way too many people supported Y. Kilcher's antics.

What antics are you referring to exactly? That he called out 'ai ethicists' who make arguments along the lines of "neural networks are bad because they cause co2 increase which hits marginalized/poor people"?

Post reply on HN