Live data from Hacker News

PixelNN – Example-Based Image Synthesis

cs.cmu.edu

51–60 of 155 posts

Re: PixelNN – Example-Based Image Synthesis

#51
post #3

I used to roll my eyes at crime television shows, whenever they said "Enhance" for a low quality image. Now it seems the possibility of that becoming realistic are increasing with a steady clip, based on this paper and other enhancement techniques I've seen posted here.

Yes!

Although what we don't have is any certainty that the enhanced face actually looks like the killer.

Re: PixelNN – Example-Based Image Synthesis

#52
post #36

Earlier quoted context omitted.

What you can do though, in limited circumstances, is create a still picture with more detail from a lower quality video. https://photo.stackexchange.com/questions/17098/csi-image-re...

It's a well-known technique in astronomy, eg https://www.aanda.org/articles/aa/ps/2005/22/aa2320-04.ps.gz

For anyone put off by the .ps.gz, it's actually just a normal web page that links to the full article in HTML and PDF. Not sure what they were thinking with that URL. I almost didn't bother to look. (Maybe that's what they were thinking?)

Re: PixelNN – Example-Based Image Synthesis

#53
post #48
post #26

Earlier quoted context omitted.

That's what video compression does now.

No, today's compression is about compressing what's already in the one movie. But imagine that you run your training set over 100's or 1000's of films, and extract just enough to represent say different types of trees in a few bytes. You could 'compress' a film by replacing data with markers that essentially describe some properties of the tree, and those properties + the training set are then used during 'decompress…

And some action movies could then be compressed to mere bytes if you basically have a virtual movie studio in your PC.

Re: PixelNN – Example-Based Image Synthesis

#54
post #30
post #7

So is there an analagous process that would apply to audio I wonder?

There kind of already is audio equivalent: MIDI. It supplies low resolution timing and pitch information and it's up to synthesizer to produce audio output matching those data.

I think the interesting part would be example based audio synthesis. Could you replace a synthesizer with a neural network which, when fed examples, would allow you to generate sounds / explore some latent space between the examples.

For example an approach similar to https://gauthamzz.github.io/2017/09/23/AudioStyleTransfer/ but then using the methods described in the PixelNN paper.

Re: PixelNN – Example-Based Image Synthesis

#55
post #3

I used to roll my eyes at crime television shows, whenever they said "Enhance" for a low quality image. Now it seems the possibility of that becoming realistic are increasing with a steady clip, based on this paper and other enhancement techniques I've seen posted here.

Except, and this is really the fundamental catch, it's not so much "enhance" as it is "project a believable substitute/interpretation". You fundamentally can't get back information that has been destroyed/or never captured in the first place. What you can do is fill in the gaps/information with plausible values. I don't know whether this sounds like I'm splitting hairs, but it's really important that the general publ…

> You fundamentally can't get back information that has been destroyed/or never captured in the first place.

I love this cliché. I've seen it thousands of times, and probably written it myself a few times. We all repeat stuff like that ad nauseam, without ever thinking.

Because it's fundamentally flawed, especially in the context that it has usually been applied to, namely criticising the CSI:XYZ trope of "enhancing images".

The truth is that there is a lot more information in a low-res image than meets the eye.

Even if you can't read the letters on a license plate, it can be recovered by an algorithm. If the Empire State Building is in the background, it's likely to be a US license plate. Maybe only some letters would result in the photo's low-res pattern. If you only see part of a letter, knowing the font may allow you to rule out many letters or numbers etc...

It's similar to that guy who used Photoshop's swirl effect to hide his face, not knowing that the effect is deterministic, and can easily be undone.

The error mostly appears to be in assuming that the information has been destroyed, when in reality it's often just obscured. And Neural Nets are excellent in squeezing all the information out noisy data.

Re: PixelNN – Example-Based Image Synthesis

#56
post #3

I used to roll my eyes at crime television shows, whenever they said "Enhance" for a low quality image. Now it seems the possibility of that becoming realistic are increasing with a steady clip, based on this paper and other enhancement techniques I've seen posted here.

Except, and this is really the fundamental catch, it's not so much "enhance" as it is "project a believable substitute/interpretation". You fundamentally can't get back information that has been destroyed/or never captured in the first place. What you can do is fill in the gaps/information with plausible values. I don't know whether this sounds like I'm splitting hairs, but it's really important that the general publ…

On the other hand, this is what the brain does all the time.

Re: PixelNN – Example-Based Image Synthesis

#57

Earlier quoted context omitted.

Except, and this is really the fundamental catch, it's not so much "enhance" as it is "project a believable substitute/interpretation". You fundamentally can't get back information that has been destroyed/or never captured in the first place. What you can do is fill in the gaps/information with plausible values. I don't know whether this sounds like I'm splitting hairs, but it's really important that the general publ…

> Except, and this is really the fundamental catch, it's not so much "enhance" as it is "project a believable substitute/interpretation". I would argue that this is a form of enhancement though, and in some cases will be enough to completely reconstruct the original image. For example, if I give you a scanned PDF, and you know for a fact that it was size 12 black Ariel text on a white background, this can feasibly le…

> The catch is that you need to know that the target image comes from roughly the same distribution as the training set.

When humans think about "enhance", they imagine extracting subtle details that were not obvious from the original, which implies that they know very little about what distribution the original image comes from. If they did, they wouldn't have a need for "enhance" 99% of the time -- the remaining 1% is for artistic purposes, which this is indeed suited for.

It'll be interesting to see how society copes with the removal of the "photographs = evidence" prior.

> when enhancing celebrity images, if the model is able to figure out who is in the picture this massively decreases the set of plausible outputs.

This is an excellent insight.

Re: PixelNN – Example-Based Image Synthesis

#58
post #53
post #48

Earlier quoted context omitted.

No, today's compression is about compressing what's already in the one movie. But imagine that you run your training set over 100's or 1000's of films, and extract just enough to represent say different types of trees in a few bytes. You could 'compress' a film by replacing data with markers that essentially describe some properties of the tree, and those properties + the training set are then used during 'decompress…

And some action movies could then be compressed to mere bytes if you basically have a virtual movie studio in your PC.

Might make it easy to watch the Sweded version.

Re: PixelNN – Example-Based Image Synthesis

#59
post #3

I used to roll my eyes at crime television shows, whenever they said "Enhance" for a low quality image. Now it seems the possibility of that becoming realistic are increasing with a steady clip, based on this paper and other enhancement techniques I've seen posted here.

Out of sheer curiosity I had a go at manually enhancing the Roundhay Garden Scene by dramatically enlarging the frames, stacking them, aligning them, erasing the most blurred ones and the obvious artifacts.

It went from this:

https://media.giphy.com/media/pUf3YfamV7BV6/giphy.gif

To this:

http://img.go-here.nl/Roundhay_Garden_Scene.gif

The funniest part was that the resolution really goes up if you make 1 px into 40 and align the frames accurately (then adjust opacity to the level of blur)

The crime television thing would be possible if you have enough frames of the gangster.

Post reply on HN