> The proof relies on a curious fact about high-dimensional geometry, which is that randomly distributed points placed on the surface of a sphere are almost all a full diameter away from each other. What theorem is this referring to? Sounds like something I should already be familiar with, but I'm not.
Computer scientists prove why bigger neural networks do better
11–20 of 151 posts
Re: Computer scientists prove why bigger neural networks do better
#12Re: Computer scientists prove why bigger neural networks do better
#13I am surprised that the paper does not even cite the Lottery Ticket Hypothesis ( https://arxiv.org/abs/1803.03635 , https://eng.uber.com/deconstructing-lottery-tickets/ ). In the LTH paper (IMHO the most fundamental deep learning publication in the last few years), the number of tickets goes as layer_size^n_layers.
Re: Computer scientists prove why bigger neural networks do better
#14> Right now, we are routinely creating neural networks that have a number of parameters more than the number of training samples. This says that the books have to be rewritten. Confused by this statement. Double descent with overparameterization is exhibited in "classical settings" too and mentioned in older books. > In their new proof, the pair show that overparameterization is necessary for a network to be robust.…
I’m curious for references or citations to this. When I was going over double descent I tried to find citations like this (just in a couple places like ML/stats textbooks).
Re: Computer scientists prove why bigger neural networks do better
#15> The proof relies on a curious fact about high-dimensional geometry, which is that randomly distributed points placed on the surface of a sphere are almost all a full diameter away from each other. What theorem is this referring to? Sounds like something I should already be familiar with, but I'm not.
Not my area of expertise, but the quoted "fact" seems at best incompletely stated: surely for it to hold there must be some constraints on the number of points (likely as a function of the diameter)?
On a high dimensional sphere they should generally be close to square root of 2 radius away from each other.
Re: Computer scientists prove why bigger neural networks do better
#16This looks interesting, I bookmarked it. My biggest blocker is the "statistics" part of M/L, knowing what algorithms to choose for various cases.
Re: Computer scientists prove why bigger neural networks do better
#17Re: Computer scientists prove why bigger neural networks do better
#18Without knowing anything about this in particular, this seems to be a rather pertinent restriction of the result related to things like sampling assumptions and the like.
Re: Computer scientists prove why bigger neural networks do better
#19> The proof relies on a curious fact about high-dimensional geometry, which is that randomly distributed points placed on the surface of a sphere are almost all a full diameter away from each other. What theorem is this referring to? Sounds like something I should already be familiar with, but I'm not.
Not my area of expertise, but the quoted "fact" seems at best incompletely stated: surely for it to hold there must be some constraints on the number of points (likely as a function of the diameter)?
Re: Computer scientists prove why bigger neural networks do better
#20> The proof relies on a curious fact about high-dimensional geometry, which is that randomly distributed points placed on the surface of a sphere are almost all a full diameter away from each other. What theorem is this referring to? Sounds like something I should already be familiar with, but I'm not.
[1] https://en.wikipedia.org/wiki/Curse_of_dimensionality
[2] http://kops.uni-konstanz.de/bitstream/handle/123456789/5715/...