Based on the numbers in the paper this is just a little bit too slow for use as a real time video effect. At ~0.1 seconds per frame we just need about a 3x improvement in performance to get to 30fps “real time” video frame rates. And on that thought since it appears they used nVidia hardware based on the CUDA dependency, it would be interesting to see how this performs on something like an M1/M2 where there’s dedicat…
nVidia hardware has dedicated ML silicon, though?
VToonify: Controllable high-resolution portrait video style transfer
11–20 of 23 posts
Re: VToonify: Controllable high-resolution portrait video style transfer
#12Based on the numbers in the paper this is just a little bit too slow for use as a real time video effect. At ~0.1 seconds per frame we just need about a 3x improvement in performance to get to 30fps “real time” video frame rates. And on that thought since it appears they used nVidia hardware based on the CUDA dependency, it would be interesting to see how this performs on something like an M1/M2 where there’s dedicat…
The paper also says they used 8 Tesla V100. They are GPUs that are dedicated for ML and quite a bit more powerful than a m2.
Re: VToonify: Controllable high-resolution portrait video style transfer
#13Based on the numbers in the paper this is just a little bit too slow for use as a real time video effect. At ~0.1 seconds per frame we just need about a 3x improvement in performance to get to 30fps “real time” video frame rates. And on that thought since it appears they used nVidia hardware based on the CUDA dependency, it would be interesting to see how this performs on something like an M1/M2 where there’s dedicat…
Does M1/M2 really outperform CUDA on beefy ML GPUs in tasks like this? I'd love to see numbers if so; this seems extremely surprising.
Re: VToonify: Controllable high-resolution portrait video style transfer
#14Based on the numbers in the paper this is just a little bit too slow for use as a real time video effect. At ~0.1 seconds per frame we just need about a 3x improvement in performance to get to 30fps “real time” video frame rates. And on that thought since it appears they used nVidia hardware based on the CUDA dependency, it would be interesting to see how this performs on something like an M1/M2 where there’s dedicat…
Does M1/M2 really outperform CUDA on beefy ML GPUs in tasks like this? I'd love to see numbers if so; this seems extremely surprising.
CUDA always requires sending data over the PCI bus, at least when it comes to realtime camera processing. GPUDirect exists but it's optimized for disks and NICs, I don't believe it's possible to use it with cameras.
Re: VToonify: Controllable high-resolution portrait video style transfer
#15Re: VToonify: Controllable high-resolution portrait video style transfer
#16Well it's cool and all, but the results sit right at the very deepest part of the uncanny valley.
It's like a tech demo, a preview of the future. Give it 5 years and it will be super refined and probably the future of low cost animation for kids TV shows and stuff. Then even further, like how no one animates without a computer now, no one will animate without AI assistance.
Re: VToonify: Controllable high-resolution portrait video style transfer
#17Earlier quoted context omitted.
The paper also says they used 8 Tesla V100. They are GPUs that are dedicated for ML and quite a bit more powerful than a m2.
Can you confirm that was for inference? I thought that was only for training 55min on 8x v100
Re: VToonify: Controllable high-resolution portrait video style transfer
#18Earlier quoted context omitted.
The paper also says they used 8 Tesla V100. They are GPUs that are dedicated for ML and quite a bit more powerful than a m2.
I missed the bit about using 8 of them to run it! Wow that’s a lot of GPU horsepower to do this. More efficient to just use a vtuber style pipeline using unreal engine and the metahumans or other avatars… only need one good GOU for that.
Re: VToonify: Controllable high-resolution portrait video style transfer
#19I wonder what the implementation into e.g live streaming would require.
Effectively these let an app (eg some VToonify tool) generate content that from the perspective of your live streaming app look like they are from a webcam
Re: VToonify: Controllable high-resolution portrait video style transfer
#20Well it's cool and all, but the results sit right at the very deepest part of the uncanny valley.
When the insensity of the style transfer is pushed mostly to the right (high), it just seems like Pixar or cartoons. Nothing uncanny whatsoever.
But when they show is about a quarter of the way to the right... it's utter nightmare fuel, like plastic surgery taken way too far. The worst kind of uncanny valley, so I definitely agree with you there.