Live data from Hacker News

Imagen Video: high definition video generation with diffusion models

imagen.research.google

31–40 of 500 posts

Re: Imagen Video: high definition video generation with diffusion models

#31
post #10

We're about a week into text-to-video models and they're already this impressive. Insane to imagine what the future holds in this space.

Insane, terrifying, incredible, etc. We're rapidly stumbling into the future of media. Who would've imagined a year ago that trivial AI image generation would not only be this advanced, but also this pervasive in the mainstream. And now video is already this good. We'll have full audio/video clips within a month.

Audio is the next thing that Stability AI is dropping, then video. In a few months you'll be able to conjure up anything you want if you have a few GPU cores. Pretty incredible.

Re: Imagen Video: high definition video generation with diffusion models

#33
The ethical implications of this are huge. Paper does a good detailing of this. Very happy to see that the researchers are being cautious.

edit: Just because it is cool to hate on AI ethics doesn't diminish the importance of using AI responsibly.

Re: Imagen Video: high definition video generation with diffusion models

#36
What everyone is missing is that these AI image/video generators lack _taste_. These tools just regurgitate a mishmash of images from it's training set, without any "feeling". What you're going to tell me that you can train them to have feeling? It's never going to happen.

Re: Imagen Video: high definition video generation with diffusion models

#37
post #29

I’m going to post an Ask HN about what am I supposed to do when I’m “disrupted”. I work in film / video / CG where the bread and butter is short form advertising for Youtube, Instagram and TV. It’s painfully obvious that in 1 year the job might be exceedingly more difficult than it is now.

When you animate a horse, does it have 5 legs with weird backwards joints? If not, your job is probably safe for now.

How long do you think until the horse looks perfect? 12 months? 5 years? I’m still 30 and I don’t see how my industry won’t be entirely disrupted by this within the next decade.

And that’s my optimistic projection. It could be we have amazing output in 24 months.

Re: Imagen Video: high definition video generation with diffusion models

#38
post #18

Someone can explains the tech limitation of the size ( 512*512 ) for those AI generated arts?

byte alignment has always been a consideration for high performance computing. this alludes to a fascinating, yet elementary, fact about computer science to me: there’s a physical atomic constraint in every algorithm.

that's not byte alignment, though- those constraints are what can be held in GPU RAM during a training batch, which is subject to a number of limits, such as "optimal texture size is a power of 2 or the next power of 2 larger than your preferred size".

Byte alignment would be more like "it's three channels of data, but we use 4 bytes (wasting 1 byte) to keep the data aligned on a platform that only allows word-level access"

Re: Imagen Video: high definition video generation with diffusion models

#40
post #20

I’m going to post an Ask HN about what am I supposed to do when I’m “disrupted”. I work in film / video / CG where the bread and butter is short form advertising for Youtube, Instagram and TV. It’s painfully obvious that in 1 year the job might be exceedingly more difficult than it is now.

Whatever insights and expertize you've gained up until now can probably be used to gain enough of a competitive advantage in this future industry to be employed. I doubt the people that will spend their time on this professionally will be former coders etc. (I've seen the stable diffusion outputs that coders will tweet. It's a good illustration that taste is still hugely important.)

I like your optimism but OP's job is to take text instructions and turn them into video, for advertisements. If Google (who already control so much of the advertising space) can take text instructions and turn them into advertisements, what's left for OP to do here? Even if there's some additional editing required this seems like it will greatly reduce the hours an editor is needed. And it can probably iterate options and work faster than a human.
Post reply on HN