Live data from Hacker News

Nvidia Vid2vid: High-resolution photorealistic video-to-video translation

github.com

21–30 of 129 posts

Re: Nvidia Vid2vid: High-resolution photorealistic video-to-video translation

#22
Some applications for this kind of tech:

- Porn, yeah, first application you can think of, there are already some startups doing it.

- Doubling actors, and applied to sound, maybe you could translate from one language to another but kind of keeping the accent and tone.

- Propaganda and misinformation. Now you can get your enemy to say and do whatever you want, on video.

- Photo-realistic games. Create a rough 3D model of an scenario and train the AI for it. Instead of photo-realistic rendering with math, render it with the AI based on a rough render, in real-time.

Re: Nvidia Vid2vid: High-resolution photorealistic video-to-video translation

#23
What media would someone collect now to be used in the future to reproduce the likeness of loved ones? Video clips of them moving? Talking? Different poses of pictures? Reading the dictionary out loud to get vocal patterns?

Heck with impersonating the POTUS. What about a lost friend, sibling or parent?

Re: Nvidia Vid2vid: High-resolution photorealistic video-to-video translation

#24

I'm probably being captain obvious here, but if this is what's being released for free, I wonder how much better a polished commercial version does, and when we reach the point where we can't trust anything we see anymore. It doesn't even have to be super perfect, even reaching the point where it takes experts about two weeks to determine if something's real or not might already be long enough to do great damage. Fro…

It could also lead to crimes like Blackmail becoming extinct. It would be hard to hold incriminating recordings of anyone over them if near-perfect audio and video synthesis was common. Especially for public figures with lots training data available.

That's equally horrifying in a different way.

Re: Nvidia Vid2vid: High-resolution photorealistic video-to-video translation

#25

Earlier quoted context omitted.

If you compare OSS or free software to commercial software generally, I don’t think there are that many massive gaps. It’s mostly polish and small incremental improvements, but the underlying tech is mostly the same. Why would that be different in this case?

Adobe have definitely bested most oss competitors in their space. Although with the amount of man power at their disposal it would be hard to beat.

The vast bulk of Adobe's advantage is in UX, not technical algorithms. Which makes perfect sense because that tends to be the case with most F/OSS software—technically brilliant but with an face only a programmer could love.

Yes, Adobe do have some remarkable algorithms that would be difficult to replicate (e.g. heal brush and content aware fill) but these are a small minority of Adobe's software advantage.

The one that irritates me the most is vector drawing programs: open source programs (and even paid competitors) just can't touch Adobe Illustrator for the sort of work I do. I'm sure at least 50 percent of it is familiarity and muscle memory, but I've desperately tried switching to a few different options like Inkscape or Affinity and left wildly disappointed.

Re: Nvidia Vid2vid: High-resolution photorealistic video-to-video translation

#26

Earlier quoted context omitted.

It could also lead to crimes like Blackmail becoming extinct. It would be hard to hold incriminating recordings of anyone over them if near-perfect audio and video synthesis was common. Especially for public figures with lots training data available.

That's equally horrifying in a different way.

Also, map images pulled from facebook to the bodies of pornstars. The creepiness and invasion of ... personal image(?) this enables is horrifying.

Re: Nvidia Vid2vid: High-resolution photorealistic video-to-video translation

#27

Earlier quoted context omitted.

Well I’m this case the results are pretty good locally but have pretty obvious artefacts too. Especially the synthesised road videos, look at the trees or even more at the Lane change in the linked video.

Can't you easily obfuscate that by making the video intentionally grainy and low-res and passing it off as "caught by CCTV" or "found footage"?

Probably much simpler to get a lookalike, and that's been possible for a long time.

Re: Nvidia Vid2vid: High-resolution photorealistic video-to-video translation

#28
post #16

I'm curious what would happen if somebody tried to impersonate the US president using this. He would say it's fake, but who would believe him? How exactly can computer scientists explain deepfakes to laymen?

Well I’m this case the results are pretty good locally but have pretty obvious artefacts too. Especially the synthesised road videos, look at the trees or even more at the Lane change in the linked video.

You say 'pretty obvious artifacts' but that is because you have some clue about the process and know what you are looking for and are interested enough to look.

I tried pointing out to some friends some really bad artifacts in a video we watched, and they just could not grasp it. They couldn't see what I was seeing as it didn't look out of place to them. It isn't for lack of intelligence, they just didn't care enough to understand. That pretty much describes vast swathes of the population.

You show a video using the above technique to anyone with strongly held political/ideologic beliefs and an inclination to accept 'alternative facts' over actual facts and videos using these techniques will be like a wildfire and almost impossible to refute!

Re: Nvidia Vid2vid: High-resolution photorealistic video-to-video translation

#29

Earlier quoted context omitted.

If you compare OSS or free software to commercial software generally, I don’t think there are that many massive gaps. It’s mostly polish and small incremental improvements, but the underlying tech is mostly the same. Why would that be different in this case?

On average I agree they are on par. Sometimes OSS is better (ffmpeg), sometimes commercial stuff (adobe after effects). In this specific case I think it might be beneficial if you can spend a lot of money on gathering training material and tweaking the network. But then again someone here mentioned the quality 4chans fake porn has reached, so maybe I'm wrong after all.

This might actually be one of those edge cases where commercial is better, but the government's secret version is by far the best.

Re: Nvidia Vid2vid: High-resolution photorealistic video-to-video translation

#30

Earlier quoted context omitted.

That's equally horrifying in a different way.

Also, map images pulled from facebook to the bodies of pornstars. The creepiness and invasion of ... personal image(?) this enables is horrifying.

That's already happened. Search for deep fakes. We already have face substitution in videos which is working surprisingly well sometimes.

There's been r/deepfakes where around 30% of the content was porn with swapped faces. It was banned though.

Post reply on HN