State-space models can learn in-context by gradient descent
1–10 of 60 posts
Re: State-space models can learn in-context by gradient descent
#2Re: State-space models can learn in-context by gradient descent
#3I don't understand. The benefit of SSMs is better scalability than self-attention. Now this adds self-attention back?
Re: State-space models can learn in-context by gradient descent
#4> We show that SSMs with local self-attention, a form of input-dependent input processing, can perform in-context learning analogously to transformers, i.e. through gradient descent steps on an implicit linear regression problem. I don't understand. The benefit of SSMs is better scalability than self-attention. Now this adds self-attention back?
Re: State-space models can learn in-context by gradient descent
#5Is there a non-autoregressive future?
Re: State-space models can learn in-context by gradient descent
#6Makes you wonder if we're training LLMs the hard way. For example, if computers had been invented before Calculus, we'd have been using "Numerical Integration" (iterating the differential squares to sum up areas, etc) and "Numerical Differentiation" (ditto for calculating slopes).
So I wonder if we're simply in a pre-Calculus-like phase of NN/Perceptrons, where we haven't yet realized there's a mathematical way to "solve" a bunch of equations simultaneously and arrive at the best (or some local minima) model weights for a given NN architecture and set of training data.
From a theoretical standpoint it IS a black box problem like this where the set of training data goes in, and an array of model weights comes out. If I were to guess I'd bet there'll be some kind of "random seed" we can add as input, and for each seed we'll get a different (local minima/maxima for model weights).
But I'm not a mathematician and there may be some sort of PROOF that what I just said can definitely never be done?
Re: State-space models can learn in-context by gradient descent
#7My own mental model for what Transformers must necessarily be doing, in order to be able to compute what they compute, given:
1. the primitives they're made of (for Transformers: matmul a learned matrix; vector-add a learned bias vector; normalize; softmax)
2. what those primitives can compute over a single layer
3. the low-ish total number of layers in a Transformer model
...is that they were already effectively "state space models" in practice. So this doesn't really surprise me!
(To be explicit, my assertion is that, for a given latent space between layers N and N+1 in a Transformer model, that latent space encodes a set of state variables [think CPU registers] used by the Nth serial computation steps of an arbitrary set of learned algorithms — where these algorithms are limited to those where every computation step is possible to encode in the form of a fused-matmul-plus-vadd, such that the algorithm itself can be learned as a depthwise-extruded sequence of weights across the layers; and where the learned algorithms can and do share state variables, both as inputs and as outputs; and where these state variables are all attenuated by an activation probability [in a Transformer: attention] such that the algorithms' outputs form a pre-multiplied conditional probability of the output given the confidence of the inputs — in turn such that the same state variable can be a low-confidence output for one algorithm, and a high-confidence output for another algorithm, and the high-confidence component of the output will swamp the low-confidence output.)
Re: State-space models can learn in-context by gradient descent
#8So they're sort of reinventing the discrete-time differentiator from signal processing, but parameterized neurally?
Re: State-space models can learn in-context by gradient descent
#9> can reproduce the outputs of an implicit linear model with least squares loss after one step of gradient descent. Makes you wonder if we're training LLMs the hard way. For example, if computers had been invented before Calculus, we'd have been using "Numerical Integration" (iterating the differential squares to sum up areas, etc) and "Numerical Differentiation" (ditto for calculating slopes). So I wonder if we're s…
Re: State-space models can learn in-context by gradient descent
#10> can reproduce the outputs of an implicit linear model with least squares loss after one step of gradient descent. Makes you wonder if we're training LLMs the hard way. For example, if computers had been invented before Calculus, we'd have been using "Numerical Integration" (iterating the differential squares to sum up areas, etc) and "Numerical Differentiation" (ditto for calculating slopes). So I wonder if we're s…
NNs have complex non-convex loss functions that don't admit a closed-form solution. Even for small models, it can be shown that it's an NP-complete problem. In fact, even for linear regression (least squares), which has a closed-form solution, it can be computationally cheaper to run gradient descent since finding the closed form solution requires you to calculate and invert a large matrix (X^T X).
Maybe our only hope of doing LLM training runs in a tiny amount of time will be from Quantum Computing or even Photonic (wave-based) Computing.