Live data from Hacker News

VideoGigaGAN: Towards detail-rich video super-resolution

videogigagan.github.io

141–150 of 243 posts

Re: VideoGigaGAN: Towards detail-rich video super-resolution

#142

Earlier quoted context omitted.

The average shot length in a modern movie is around 2.5 seconds (down from 12 seconds in 1930's). For animations it's around 15 seconds.

The textures of objects need to maintain consistency across much larger time frames, especially at 4k where you can see the pores on someone's face in a closeup.

Off topic: the clarity of pores and fine facial hair on Vision Pro when watching on a virtual 120-foot screen is mindblowing.

Re: VideoGigaGAN: Towards detail-rich video super-resolution

#143
post #33

Would be neat to see this on much older videos (maybe WW2 era) to see how it improves details.

That is essentially what Peter Jackson did for the 2018 film They Shall Not Grow Old: https://www.imdb.com/title/tt7905466/

They used digital upsampling techniques and colorization to make World War One footage into high resolution. Jackson would later do the same process for the 2021 series Get Back, upscaling 16mm footage of the Beatles taken in 1969: https://www.imdb.com/title/tt9735318/

Both of these are really impressive. They look like they were shot on high resolution film recently, instead of fifty or a hundred years ago. It appears that what Peter Jackson and his team did meticulously at great effort can now be automated.

Everyone should understand the limitations of this process. It can't magically extract details from images that aren't there. It is guessing and inventing details that don't really exist. As long as everyone understands this, it shouldn't be a problem. Like, we don't care that the cross-stitch on someone's shirt in the background doesn't match reality so long as it's not an important detail. But if you try to go Blade Runner/CSI and extract faces from reflections of background objects, you're asking for trouble.

Re: VideoGigaGAN: Towards detail-rich video super-resolution

#145

Wonder how long until Hollywood CGI shops have these types of models running as part of their post-production pipeline. Big blockbusters often release with ridiculously broken CGI due to crunch (Black Panther's third act was notorious for looking like a retro video-game), adding some extra generative polish in those cases is a no-brainer.

Once AI tech gets fully integrated entire Hollywood rendering pipeline will go from rendering to diffusing

Once AI tech gets fully integrated, the movie industry will cease to exist.

Re: VideoGigaGAN: Towards detail-rich video super-resolution

#148

Video quality seems really good, but limitations are quite restrictive "Our model encounters challenges when processing extremely long videos (e.g. 200 frames or more)". I'd say most videos in practice are longer than 200 frames, so lot more research is still needed.

Wonder what happens if you run it piece-wise on every 200 frames. Perhaps it glitches in the interface.

Re: VideoGigaGAN: Towards detail-rich video super-resolution

#150

It's impressive, but still looks kinda bad? I think the video of the camera operator on the ladder shows the artifacts the best. The main camera equipment is no longer grounded in reality, with the fiddly bits disconnected from the whole and moving around. The smaller camera is barely recognizable. The plant in the background looks blurry and weird, the mountains have extra detail. Finally, the lens flare shifts! Che…

I think the hand running through the wheat (?) is pretty good, object permanence is pretty reasonable especially considering the GAN architecture. GANs are good at grounded generation--this is why the original GigaGAN paper is still in use by a number of top image labs. Inferring object permanence and object dynamics is pretty impressive for this structure.

Plus, a rather small data set: REDS and Vimeo-90k aren't massive in comparison to what people speculate Sora was trained on.

Post reply on HN