Live data from Hacker News

Imagen Video: high definition video generation with diffusion models

imagen.research.google

241–250 of 500 posts

Re: Imagen Video: high definition video generation with diffusion models

#242

Earlier quoted context omitted.

"Stable Diffusion" is a particular brand from the company Stability AI that is famously open sourcing all of their models.

Pedantically, Stable Diffusion v1.4 is the one model where weights were open sourced and released. Stable Diffusion v1.5, announced September 8th and live on their API, was to be released in "a week or two" but still has yet to be released to the general public. https://discord.com/channels/1002292111942635562/10022921127...

SD 1.2 and 1.3 are open source too

Re: Imagen Video: high definition video generation with diffusion models

#243

"We have decided not to release the Imagen Video model or its source code until these concerns are mitigated" Okay then why even post it in the first place? What exactly is Google going to do with this model?

Why post? to show methods and their capabilities. Also flex. What will they do with model? figure out how to prevent abuse and incorporate into future Google Assistant, Photos and AR offerings.

Just fixing their basic stuff would be a better start from where they are right now.

Re: Imagen Video: high definition video generation with diffusion models

#244

We're about a week into text-to-video models and they're already this impressive. Insane to imagine what the future holds in this space.

How is it possible that all of them just started to appear at the same time? Is it possible that those models were designed and trained in a last few weeks? Has some "magic key" to content generation been just unexpectedly discovered? Or the topic became trendy and everyone is just publishing what they've got so far, so they hope to benefit from media attention?

This is why

https://www.reddit.com/r/singularity/comments/xwdzr5/the_num...

Re: Imagen Video: high definition video generation with diffusion models

#245
post #43

This sort of AI related work seems to be accelerating at an insane speed recently. I remember being super impressed by AI Dungeon and now in the span of a few months we have got DALLE-2 , Stable Diffussion, Imagen, that one AI powered video editor, etc. Where do we think we will be at in 5 years??

I'd say in less than 10 years we will be able to turn novels into movies using deep learning at this rate.

Re: Imagen Video: high definition video generation with diffusion models

#246
post #124

Earlier quoted context omitted.

Try again and do what? How are you directing the shot? How do you erase an emotion? How do you erase and redo inner turmoil when delivering a performance?

You tell it, "do it all over again, now with less inner turmoil". Not joking, that's all it's going to take. There are also a few diffusion based speech generators that handle all sounds, inflections and styles, they are going to come in handy for tweaking turmoil levels.

Yep!

"Restyle that last scene, showing different mixtures of fear/concern/excitement on male lead's face. Try to evoke a little of Harrison Ford's expressions in his famous roles. Render me 20 alternate treatments."

[5 minutes later]

«Here are the 20 alternate takes you requested for ranking.»

"OK, combine take #7 up to the glance back, with #13 thereafter."

«Done.»

Re: Imagen Video: high definition video generation with diffusion models

#247

What everyone is missing is that these AI image/video generators lack _taste_. These tools just regurgitate a mishmash of images from it's training set, without any "feeling". What you're going to tell me that you can train them to have feeling? It's never going to happen.

They work at the level of convolutions, not images.

Re: Imagen Video: high definition video generation with diffusion models

#249
> We train our models on a combination of an internal dataset consisting of 14 million video-text pairs

The paper is sorely lacking evaluation; one thing I'd like to see for instance (any time a generative model is trained on such a vast corpus of data) is a baseline comparison to nearest-neighbor retrieval from the training data set.

Re: Imagen Video: high definition video generation with diffusion models

#250
post #102

These are baby steps towards what I think will be the eventual "disruption" to the film and tv industry. Directors will simply be able to write a script/prompt long enough and detailed enough for something like Imagen (or it's successors) to convert into a feature-length show. Certainly we're very, very far away from that level of cinematic detail and crispness. But I believe that is where this leads... complete with…

There's already a surplus of video and an apparent lack of _quality_ video. This might be enough to get folks to shut the TV off completely.

Has this alleged lack of quality video caused total consumption of televised entertainment to decline recently?
Post reply on HN