Mamba: The Easy Way
31–40 of 63 posts
Re: Mamba: The Easy Way
#32In case people are wondering why Mamba is exciting: There's this idea in AI right now that "scaling" models to be bigger and train on more data always makes them better. This has led to a science of "scaling laws" which study just how much bigger models need to be and how much data we need to train them on to make them a certain amount better. The relationship between model size, training data size, and performance t…
I'd love to see someone who has the resources train a model bigger than 2.8b and show the scaling law still holds.
Re: Mamba: The Easy Way
#33All of the sources I see referred to as derivations of it have a discretization of the form
h_t =Ah_{t-1} + Bx_{t-1} for the first line instead of the given one of the form h_t =Ah_{t-1} + Bx_t.
Does anyone know why this is?
Re: Mamba: The Easy Way
#34In case people are wondering why Mamba is exciting: There's this idea in AI right now that "scaling" models to be bigger and train on more data always makes them better. This has led to a science of "scaling laws" which study just how much bigger models need to be and how much data we need to train them on to make them a certain amount better. The relationship between model size, training data size, and performance t…
Re: Mamba: The Easy Way
#35From what I can tell all the large players in the space are continuing developing on transformers right? Is it just that Mamba is too new, or is the architecture fundamentally not usable for some reason?
we have no idea what the large players in the space are doing
Re: Mamba: The Easy Way
#36I'm very positive I can actually understand the terminology used in discussing machine learning models if it was presented in a way that describes the first principles a little bit better, instead of diving directly into high level abstract equations and symbols. I'd like a way to learn this stuff as a computer engineer, in the same spirit as "big scary math symbols are just for-loops"
Ironically, you can probably just ask a Transformer model to explain it to you. I'm the same as you: I have no problem grasping complex concepts, I just always struggled with the mathematical notation. I did pass linear algebra in university, but was glad I could go back to programming after that. Even then, I mostly passed linear algebra because I wrote functions that solve linear algebra equations until I fully gra…
Also, just try really hard. Repeat. It's new language to explain concepts you likely already know. You don't remember spanish by looking at the translations once.
Re: Mamba: The Easy Way
#37This was really helpful, but only discusses linear operations, which obviously can’t be the whole story. From the paper it seems like the discretization is the only nonlinear step—in particular the selection mechanism is just a linear transformation. Is that right? How important is the particular form of the nonlinearity? EDIT: from looking at the paper, it seems like even though the core state space model/selection…
Re: Mamba: The Easy Way
#38If I'm not mistaken the largest mamba model right now is 2.8B and undertrained with low quality data (the Pile only). The main problem is that it's new and unproven. Should become very interesting once someone with both data and significant financial backing takes the plunge and trains something of notable size. Perhaps Llama-3 might already end up being that attempt, as we seem to be heavily into diminishing returns…
There is one trained on 600B tokens from SlimPajama [1], but that's fairly tiny compared to other recent releases (ex. stablelm-3b [2] trained on 4T tokens). > low quality data (the Pile only) The Pile is pretty good quality wise. It's mostly the size (300B tokens) that's limiting. [1]: https://huggingface.co/state-spaces/mamba-2.8b-slimpj [2]: https://huggingface.co/stabilityai/stablelm-3b-4e1t
Re: Mamba: The Easy Way
#39Fantastic blog post, thank you for this. I am not even familiar with transformers, yet the explanation is stellar clear to me, and the included references and context are a trasure trove. The explanation of FlashAttention is the best I have seen, and that is not even the focus of the article. One question I have on selectivity: footnote 4 says "the continuous A is constant, while our discretization parameter ∆ is inp…
How are you not familiar with transformers yet have seen multiple explanations of FlashAttention?
Re: Mamba: The Easy Way
#40Something minor I always wonder about when I read Mamba is the discretization. All of the sources I see referred to as derivations of it have a discretization of the form h_t =Ah_{t-1} + Bx_{t-1} for the first line instead of the given one of the form h_t =Ah_{t-1} + Bx_t. Does anyone know why this is?