Live data from Hacker News

Deep image prior 'learns' on just one image

dmitryulyanov.github.io

111–120 of 235 posts

Re: Deep image prior 'learns' on just one image

#111

Earlier quoted context omitted.

There is some pretty strong evidence for this: all the toddlers in the world. You only need to show them something once and they'll immediately be able to recognize more examples of the same thing from different angles and even when it is partially hidden. All they have to guide them is the structure of their brains, not the quantity of data they have been exposed.

Seems like this view gets told every once in a while by someone who clearly hasn't been around any 0-2 year olds.

This is HN, what did you expect? Real people with wife and kids?

Re: Deep image prior 'learns' on just one image

#113
As I'm not an expert in the field; what exactly does the term

  min_x E(x; x0) + R(x)
mean?

I thought that E(x; x0) would denote the error/difference between original and corrupted images, and R(x) be the (searched-for) correction. But this doesn't seem to make sense with the next parts of their explanation.

Re: Deep image prior 'learns' on just one image

#114
post #21

Wow: "In this work, we show that, contrary to expectations, a great deal of image statistics are captured by the structure of a convolutional image generator rather than by any learned capability. This is particularly true for the statistics required to solve various image restoration problems, where the image prior is required to integrate information lost in the degradation processes. To show this, we apply untrain…

So are they saying that the topology of a deep-net is intrinsic to "reality" ... somewhat analogous to something like the Fibonacci ratio for organic forms?

The Fibonacci thing is mostly a myth perpetuated by confirmation bias tough.

Re: Deep image prior 'learns' on just one image

#115

I'm finding it hard to put into words what I find wrong with this paper, but ... here goes nothing. So, the novel thing here is that an encoder-decoder network applied to an image can learn enough from a source image to be useful. In some ways that's obvious, but the effectiveness of it on reconstruction tasks is certainly surprising. I have two problems, though. One is that I would take the reconstruction results wi…

>Another example from the annals of machine learning history: time and time again when there is a breakthrough in architectures, it's usually followed a few years later by a simplification of the architecture. For example we started with networks like VGG which are these big, hand crafted architectures. Slowly over time architectures have become _less_ exotic, instead opting to simply define a basic building block repeated N times.

>The reason for this is because in the intervening years we gather more training data, better training techniques, and more computational power. So we can instead use a more homogeneous architecture which has _less_ assumptions (less priors) and at the end of the day we get _better_ results.

>I'll repeat that. We put _less_ priors into our networks and we get better results.

Is move from VGG-like architectures a question of moving to using less priors or just move to different priors that work better?

In addition, I'd claim that a choice of training method also imbues the whole ML system with prior knowledge. If we add a dropout layer, we wish to enforce a particular regularization scheme on the model because we know that we need to regularize to model to avoid overfitting. If we choose to use a particular optimizer instead of another, we use it because we know it has desirable properties on particular objective objection.

>So we already know that architecture isn't ultimately important. You can have a giant, fully connected network, and it _will_ work, if you feed it enough data and have the computational power necessary to train such a beast.

I'm slightly unclear about your argument here. What is the point of training the fully connected network for a long time and with massive datasets to get equivalent results and structure as convnets if we already do know that convolutional structure with weight sharing is a good fit for natural images? Deep learning has been casted as a magical black box that produces impressive results "yet the scientists don't know why!" by the press. On the contrary, these results suggest there might be less magic in neural net models as the recent hype would suggest.

> The examples are clearly lab queens, where the occluded regions are not particularly interesting/challenging.

To me they look quite like the standard examples used in inpainting and denoising articles.

Re: Deep image prior 'learns' on just one image

#116
post #90
post #21

Wow: "In this work, we show that, contrary to expectations, a great deal of image statistics are captured by the structure of a convolutional image generator rather than by any learned capability. This is particularly true for the statistics required to solve various image restoration problems, where the image prior is required to integrate information lost in the degradation processes. To show this, we apply untrain…

I don't understand their editorializing > contrary to expectations, ingore the weasel words > a great deal of image statistics are captured by the structure of a convolutional image generator rather than by any learned capability. > To show this, we apply untrained ConvNets to the solution of several such problems. Instead of following the common paradigm of training a ConvNet on a large dataset of example images, we…

Agreed. The network is trained in a sense "online" on the target image and directly applied to it. I don't see much impact of this work outside image denouncing or tasks already presented in the paper.

Re: Deep image prior 'learns' on just one image

#117
post #21

Wow: "In this work, we show that, contrary to expectations, a great deal of image statistics are captured by the structure of a convolutional image generator rather than by any learned capability. This is particularly true for the statistics required to solve various image restoration problems, where the image prior is required to integrate information lost in the degradation processes. To show this, we apply untrain…

Does this also explain - on a higher level - why AlphaGo Zero was possible and so successful?

Re: Deep image prior 'learns' on just one image

#118
post #117
post #21

Wow: "In this work, we show that, contrary to expectations, a great deal of image statistics are captured by the structure of a convolutional image generator rather than by any learned capability. This is particularly true for the statistics required to solve various image restoration problems, where the image prior is required to integrate information lost in the degradation processes. To show this, we apply untrain…

Does this also explain - on a higher level - why AlphaGo Zero was possible and so successful?

Alpha Go Zero is actually more similar to a GAN: https://arxiv.org/abs/1711.09091

Re: Deep image prior 'learns' on just one image

#119
post #32
post #30

Earlier quoted context omitted.

> PS. This makes me wonder whether and to what degree the structure of the brain's connectome is a necessary prior for AGI. Well, I wouldn't mix up AGI and AGI by deep learning, and more important I would emphasise that this is a good prior for images . The fundamental insight in CNNs and eventually in this work is that there is a correlation between pairs of nearby pixels. We have something similar for video and aud…

Thanks. I'm not mixing them up! I'm just wondering whether and to what degree architecture , i.e., network structure, will prove important for other, more advanced AI tasks, including up to AGI.

To throw a dissenting voice into the mix:

I do not think intelligence is a consequence of structure. Transparency is a consequence of how light interacts with objects. There is no "transparent gold". And being transparent is not something we can program gold to do.

Programming is the application of an electric field across a silicon surface: this can no more transmute silicon into gold as it can into a nervous system.

Consciousness is a biological activity. It is not something wood, silicon, sand, metal, glass, water, etc. can do by virtue of moving around in an interesting pattern. To me, this is akin to believing magic, or to thinking that the issue with Bug's Bunny's consciousness is that the ink wasnt laid out in the right way. Rather, there isnt any.

As scientists we should begin very clearly by identifying the target phenomenon. Line up everything in the known universe that is conscious and you will find that it engages in a highly specific biochemical activity that requires a whole set of highly specific chemical interactions.

To believe these can be achieved by applying an electric field across some silicon seems like the victorian superstition that a person could be raised from the dead by doing likewise.

Rather, all we are able to do is imitate in the weakest sense. People fooled by chatbots are no more enlightened than the early victorians who were shocked to watch the movie of a train coming at them. Fooling people is a very easy thing to do. And esp. in the area of consciousness when "alertness to things that might be conscious" seems a deep psychological preoccupation of almost every animal to the point of absurdity (eg. being scared at the wind).

Re: Deep image prior 'learns' on just one image

#120

As I'm not an expert in the field; what exactly does the term min_x E(x; x0) + R(x) mean? I thought that E(x; x0) would denote the error/difference between original and corrupted images, and R(x) be the (searched-for) correction. But this doesn't seem to make sense with the next parts of their explanation.

x is actually the generated image they are testing against x0. A lower E(x; x0) means an image which fits well towards the objective based on the original image (depends on the task). The paper gives some examples. For example, for the task of image denoising, E(x; x0) is just the squared distance of the generated (denoised candidate) image x to the pixels to the original image x0. Obviously you would want this to be low in the generated version since it should still look close to x0.

R(x) is a regularization term to avoid overfitting. For example, in the denoising example, it could be a measure of the variation in color of x. Clearly, just taking x = x0, the squared error (E(x; x0)) is 0, but it will have high R(x) because of all the noise. That's why they try to minimize both quantities combined, so we get min_x E(x; x0) + R(x), to get close to the objective but also not overfit.

https://en.wikipedia.org/wiki/Regularization_(mathematics)

Post reply on HN