Live data from Hacker News

Stacked Approximated Regression Machine: A Simple Deep Learning Approach

arxiv.org

21–28 of 28 posts

Re: Stacked Approximated Regression Machine: A Simple Deep Learning Approach

#21
post #10

Earlier quoted context omitted.

One of the authors, Zhangyang Wang, just wrote this on his personal page: "We have discussed and decided to work on a software package release, perhaps accompanying it with a more detailed technical report in the future. Once the software package is ready, we will update everybody." http://www.atlaswang.com/

Paper withdrawn https://arxiv.org/abs/1608.04062 It's kind of an odd thing. I (random non-academic amateur) actually spent a bunch of time trying to parse the paper, which was kind of a combination of interesting ideas and incomprehensible ambiguities. One real academic researcher also put some time into it. The good part of the paper is explained here. My guess is the problem is going from ARM to SARM. http://gabgoh…

Hi, I'm the author of the blog post. I added a blurb to the beginning the blog post explaining all the drama, and precisely what claim was made that was withdrawn.

The problem is not in the ARG->SARG approximation, but the bit on unsupervised pretraining. The paper could stand on its own without that section, but without that result it would have been a significantly more mediocre NIPS submission. Hope this clarifies things.

Re: Stacked Approximated Regression Machine: A Simple Deep Learning Approach

#22

Earlier quoted context omitted.

Paper withdrawn https://arxiv.org/abs/1608.04062 It's kind of an odd thing. I (random non-academic amateur) actually spent a bunch of time trying to parse the paper, which was kind of a combination of interesting ideas and incomprehensible ambiguities. One real academic researcher also put some time into it. The good part of the paper is explained here. My guess is the problem is going from ARM to SARM. http://gabgoh…

Hi, I'm the author of the blog post. I added a blurb to the beginning the blog post explaining all the drama, and precisely what claim was made that was withdrawn. The problem is not in the ARG->SARG approximation, but the bit on unsupervised pretraining. The paper could stand on its own without that section, but without that result it would have been a significantly more mediocre NIPS submission. Hope this clarifies…

Your blog post was excellent, btw.

Re: Stacked Approximated Regression Machine: A Simple Deep Learning Approach

#23
post #13

Earlier quoted context omitted.

This is a pretty good summary, but to note this is not a new unsupervised technique - it's a sparse coding method, and sparse coding has been around for more than a decade. Summarizing this paper without mentioning the word sparse coding seems wrong somehow. ;)

You're right, sparse coding is not new -- and neither is the idea of greedy layer-wise training. But if this works as advertised, these guys must be doing something new .

But, sadly, it doesn't work as advertised. My hypothesis is that it's extremely difficult to get great results with greedy training. Jointly learning multiple layers from huge datasets is why things started working so well.

Re: Stacked Approximated Regression Machine: A Simple Deep Learning Approach

#24

Earlier quoted context omitted.

Paper withdrawn https://arxiv.org/abs/1608.04062 It's kind of an odd thing. I (random non-academic amateur) actually spent a bunch of time trying to parse the paper, which was kind of a combination of interesting ideas and incomprehensible ambiguities. One real academic researcher also put some time into it. The good part of the paper is explained here. My guess is the problem is going from ARM to SARM. http://gabgoh…

Hi, I'm the author of the blog post. I added a blurb to the beginning the blog post explaining all the drama, and precisely what claim was made that was withdrawn. The problem is not in the ARG->SARG approximation, but the bit on unsupervised pretraining. The paper could stand on its own without that section, but without that result it would have been a significantly more mediocre NIPS submission. Hope this clarifies…

First, thanks for the excellent blog, it gave me a better idea what was happening

As far as the ARG-> transformation goes, maybe that's just something I don't get, I can see how one goes from sparse encoding to repeated ARG-type transformations and how this repeated application approximates the solution of a sparse encoding problem. And it is suggestive that these application look like a layers of a neural net.

But when you switch to stacking, what are you doing? Solving one sparse encoding problem then another? What analogy is there to say this works ... or that it would work better than just single sparse encoding? At that point, is it just "try it and see?"

One of the impressions I got from scanning the literature is that deep nets are kind of generally treacherous beasts - just getting a locally 1st layer may not be desirable off the bat. People have settled on backpropagation for very subtle reasons. See "Overfitting in Neural Nets: Backpropagation, Conjugate Gradient, and Early Stopping", Caruana, Lawrence, et. al where backpropagation finds better solutions than the "more powerful" conjugate gradient method.

Re: Stacked Approximated Regression Machine: A Simple Deep Learning Approach

#25
post #23
post #13

Earlier quoted context omitted.

You're right, sparse coding is not new -- and neither is the idea of greedy layer-wise training. But if this works as advertised, these guys must be doing something new .

But, sadly, it doesn't work as advertised. My hypothesis is that it's extremely difficult to get great results with greedy training. Jointly learning multiple layers from huge datasets is why things started working so well.

Yes, I just saw the retraction.

Re: Stacked Approximated Regression Machine: A Simple Deep Learning Approach

#26

Earlier quoted context omitted.

Paper withdrawn https://arxiv.org/abs/1608.04062 It's kind of an odd thing. I (random non-academic amateur) actually spent a bunch of time trying to parse the paper, which was kind of a combination of interesting ideas and incomprehensible ambiguities. One real academic researcher also put some time into it. The good part of the paper is explained here. My guess is the problem is going from ARM to SARM. http://gabgoh…

Hi, I'm the author of the blog post. I added a blurb to the beginning the blog post explaining all the drama, and precisely what claim was made that was withdrawn. The problem is not in the ARG->SARG approximation, but the bit on unsupervised pretraining. The paper could stand on its own without that section, but without that result it would have been a significantly more mediocre NIPS submission. Hope this clarifies…

gabrielgoh: thank you for this. Like others, I was more optimistic than others that this had promise, and I was wrong. Now I want to understand how the authors' got it wrong, so I've added your blog post to my reading list.

Re: Stacked Approximated Regression Machine: A Simple Deep Learning Approach

#27

Earlier quoted context omitted.

Hi, I'm the author of the blog post. I added a blurb to the beginning the blog post explaining all the drama, and precisely what claim was made that was withdrawn. The problem is not in the ARG->SARG approximation, but the bit on unsupervised pretraining. The paper could stand on its own without that section, but without that result it would have been a significantly more mediocre NIPS submission. Hope this clarifies…

First, thanks for the excellent blog, it gave me a better idea what was happening As far as the ARG-> transformation goes, maybe that's just something I don't get, I can see how one goes from sparse encoding to repeated ARG-type transformations and how this repeated application approximates the solution of a sparse encoding problem. And it is suggestive that these application look like a layers of a neural net. But w…

you are right. you are using the output of the previous sparse solution as input into the new one, i.e. stacking sparse coders.

Your second question of why this is a good idea is the million dollar question. Its pretty much "lets try it and see", with some heuristic reasoning thrown into the mix (its mirrors the brain, it abstracts information, etc, etc).

btw, I don't think people use early stopping anymore. It's been replaced by more powerful forms of regularization, such as dropout. The deep learning world is getting more tame, and that makes me happy.

Re: Stacked Approximated Regression Machine: A Simple Deep Learning Approach

#28
post #2

As far as I understand this, these guys claim they can train convolutional and many other types of deep neural nets faster by pretraining each layer with a new unsupervised technique via which the layer sort of learns to compress its inputs (a local optimization problem), and then they fine tune the whole network end-to-end with supervised SGD and backpropagation as usual. They have not released code, so no one else…

The paper itself is fairly sparse but references a number of approaches for more-quickly-learning neural-net-related learning systems from ~2013 (PCAnet, SCATnet,etc). The paper presents these and other approaches as being instances of a classical, general form, regularized regression but with the "stacked" property involving each layer iterating however many times and then the next layer changing parameters (or feat…

Note: even though the paper talks about Approximate Regression Machine layers and uses equations that look sort-of like the equations of regularized regression, the layers aren't about regularized regression but about sparse dictionary encoding, a quite different approach.
Post reply on HN