Live data from Hacker News

Imagen Video: high definition video generation with diffusion models

imagen.research.google

141–150 of 500 posts

Re: Imagen Video: high definition video generation with diffusion models

#141

I’m going to post an Ask HN about what am I supposed to do when I’m “disrupted”. I work in film / video / CG where the bread and butter is short form advertising for Youtube, Instagram and TV. It’s painfully obvious that in 1 year the job might be exceedingly more difficult than it is now.

I think short advertisements would be affected most by this, it seems.

But here is the catch, there is the same last mile problem for those AI models. Currently it feels like the model can achieve like 80%-90% what a trained human expert can do, but the last 10-20% would extra extra hard to reach human fidelity. It might take years, or it might never happen.

That being said, I think anyone who doubts AI-assisted creative workflow is a fuzz is deadly wrong, anyone who refuses those shiny new tools, is likely to be eliminated by sheer market dynamics. They can't compete on the efficiency of it.

Re: Imagen Video: high definition video generation with diffusion models

#143

Google continues to blow my mind with these models, but I think their ethics strategy is totally misguided and will result in them failing to capture this market. The original Google Search gave similarly never-before-seen capabilities to people, and you could use it for good or bad - Google did not seem to have any ethical concerns around, for example, letting children use their product and come across NSFW content…

It makes sense though. The biggest threat to Google right now isn't some scrappy startup eating their lunch. It's the looming regulatory action over antitrust and privacy that could weaken or destroy their core business. As this is a political problem (not a technical one), they don't want to do anything that could upset politicians or turn public opinion against them. Personally, I doubt they have serious ethical concerns over releasing the model. I do believe they have serious "AI ethics 'thought leaders' and politicians will use this against us" concerns.

Re: Imagen Video: high definition video generation with diffusion models

#144

Google continues to blow my mind with these models, but I think their ethics strategy is totally misguided and will result in them failing to capture this market. The original Google Search gave similarly never-before-seen capabilities to people, and you could use it for good or bad - Google did not seem to have any ethical concerns around, for example, letting children use their product and come across NSFW content…

I agree, but I also think that the ethics is just an excuse not to release the source code & models. The AI community clearly disapproves of papers without code. This is a way to skirt around that disapproval. You get to keep the code and models private and (they hope) not be criticised for it.

With Stable Diffusion I think they just didn't expect someone to produce a truly open version. There are plenty of AI models that Google have made where they've maintained a competitive advantage for many years by not releasing the code/models, e.g. speech recognition.

Re: Imagen Video: high definition video generation with diffusion models

#145

Google continues to blow my mind with these models, but I think their ethics strategy is totally misguided and will result in them failing to capture this market. The original Google Search gave similarly never-before-seen capabilities to people, and you could use it for good or bad - Google did not seem to have any ethical concerns around, for example, letting children use their product and come across NSFW content…

Perhaps Google hasn't found the right balance in this case, but as a general rule, less ethics === more market. This isn't unique in that way.

Re: Imagen Video: high definition video generation with diffusion models

#146
post #18

Someone can explains the tech limitation of the size ( 512*512 ) for those AI generated arts?

It's limited by the RAM on the GPU, with most consumer-grade cards having closer to 8 GiB VRAM than the 80 GiB VRAM datacenter cards have.

Re: Imagen Video: high definition video generation with diffusion models

#147

Google continues to blow my mind with these models, but I think their ethics strategy is totally misguided and will result in them failing to capture this market. The original Google Search gave similarly never-before-seen capabilities to people, and you could use it for good or bad - Google did not seem to have any ethical concerns around, for example, letting children use their product and come across NSFW content…

Google is absolutely not going to start taking more risks. They are at the part of the business lifecycle where they squeeze the juice out of the cash cow and protect it jealously in the meantime. While Google gets much recognition for this research, I believe they are incapable as a corporate entity of creating a product out of it because they can no longer capable of taking risks. That is going to fall to other companies still building their product and able to gamble on risk-reward.

Re: Imagen Video: high definition video generation with diffusion models

#148
post #124
post #123

Earlier quoted context omitted.

I disagree. It's a rudimentary features of all these models to take a concept picture and refine it. It won't be like the director would give a prompt and get a feature length movie, it will be more like the director uses MS Paint (as in a common software for non tech people) to make a scene outline and directs AI to make a stylish and animated version of that. Something is wrong? just erase it and try again. Dalle2…

Try again and do what? How are you directing the shot? How do you erase an emotion? How do you erase and redo inner turmoil when delivering a performance?

You tell it, "do it all over again, now with less inner turmoil". Not joking, that's all it's going to take. There are also a few diffusion based speech generators that handle all sounds, inflections and styles, they are going to come in handy for tweaking turmoil levels.

Re: Imagen Video: high definition video generation with diffusion models

#149
post #115

I’ll be honest, as someone who worked in the film industry for a decade, this thread is depressing. It’s not the technology, it’s all the people in these comments who have never worked in the industry clamouring for its demise. One could brush it off as tech heads being over exuberant, but it’s the lack of understanding of how much fine control goes into each and every shot of a film that is depressing. If I, as a cr…

The term "creative" is so pretentious, as if only content generation involves creativity.

Your post reminds me of all the photographers that said digital photography would remain niche and never replace film.

The current models are toys made by small groups. It's not hard to imagine AI generated film being much more compelling when the entire industry of engineers and "creatives" refine and evolve the ecosystem to take into account subtle strokes, wrinkles, movement, shots etc. And they will, because it will be cheaper, and businesses always go for cheaper.

Re: Imagen Video: high definition video generation with diffusion models

#150
post #122
post #116

Earlier quoted context omitted.

There's no doubt that it's only a matter of time. Like bloggers had the opportunity to compete with newspapers, the ability to generate videos will allow to compete with movies/marvel/netflix/disney & company. Eventually, only high quality content will justify the need to pay for a ticket or a subscription, and there's going to be a lot of free content to watch, with 1000x more people able to publish their ideas, as…

You’re conflating the ability to make things for the masses and being able to automatically generate it. Film production is already commoditized and anyone can make high end content. Being able to automatically create that is a different argument than what you posit.

I don't think this matters, new movies and TV shows already have to compete with a huge amount of old content, some of it amazing. Just like a new painting or professional photo has to compete with the billions of images already existing on the web. Generative models for video and image are not going to change the fact we already can't keep up.
Post reply on HN