A pure NumPy implementation of Mamba
11–20 of 39 posts
Re: A pure NumPy implementation of Mamba
#12For me pure X means: to use this, all you have to install is X.
Re: A pure NumPy implementation of Mamba
#13Totally unclear what this is. I scrolled through the readme and it didn't even mention once what it does.
I also assumed that "a pure NumPy implementation" meant that it was built purely with numpy, which it isn't smh
Re: A pure NumPy implementation of Mamba
#14Earlier quoted context omitted.
It’s much more than just an LLM. The mamba architecture is often used in the backbone of an LLM but you can use it more generally as a linear-time (as opposed to quadratic-time) sequence modeling architecture (as per the original paper’s title, which is cited in the linked repo). It is much closer to a convolutional network or an RNN (it has bits of both) than to a transformer architecture. It is based off the notion…
Would love to hear more about that building energy modelling example, have you done a writeup you could share?
I can also share my master’s thesis which is similar but using CNN layers rather than Mamba and only for monthly predictions rather than 15-min interval data. There are some other architectural differences but the basics are the same. That work is also globally robust.
As you can imagine, the current work I am doing at a much higher resolution is a big step up, and Mamba so far is working out great.
Re: A pure NumPy implementation of Mamba
#15Is there a benefit to implementing it in numpy over pytorch or tf?
A numpy program will work tomorrow.
ALL of the machine learning frameworks have incredible churn. I have code from two years ago which I can't make work reliably anymore -- not for lack of trying -- due to all the breaking changes and dependency issues. There are systems where each model runs in its own docker, with its own set of pinned library versions (many with security issues now). It's a complete and utter trainwreck. Don't even get me started on CUDA versions (or Intel/AMD compatibility, or older / deprecated GPUs).
For comparison, virtually all of my non-machine-learning Python code from the year 2010 all still works in 2024.
There are good reasons for this. Those breaking changes aren't just for fun; they're representative of the very rapid rate of progress in a rapidly-changing field. In contrast, Python or numpy are mature systems. Still, it makes many machine learning models insanely expensive to maintain in production environments.
If you're a machine learning researcher, it's fine, but if you have a system like an ecommerce web site or a compiler or whatever, where you'd like to be able to plug in a task-specific ML model, your downpayment is a weekend of hacking to make it work, but your ongoing rent of maintenance costs might be a few weeks each year for each model you use. I have a million places I'd love to plug in a little bit of ML. However, I'm very judicious with it, not because it's hard to do, but because it's expensive to maintain.
A pure Python + numpy implementation would mean that you can avoid all of that.
Re: A pure NumPy implementation of Mamba
#16Earlier quoted context omitted.
It totally mentions what it does. It takes the sentence "I have a dream that" and extends it to: "I have a dream that I will be able to see the sunrise in the morning." It's an LLM.
It’s much more than just an LLM. The mamba architecture is often used in the backbone of an LLM but you can use it more generally as a linear-time (as opposed to quadratic-time) sequence modeling architecture (as per the original paper’s title, which is cited in the linked repo). It is much closer to a convolutional network or an RNN (it has bits of both) than to a transformer architecture. It is based off the notion…
Re: A pure NumPy implementation of Mamba
#17Totally unclear what this is. I scrolled through the readme and it didn't even mention once what it does.
It's an LLM architecture competing with transformers: https://arxiv.org/abs/2312.00752 Proponents of it usually highlight it's inference performance, in particular linear scaling with the input tokens.
Re: A pure NumPy implementation of Mamba
#18Earlier quoted context omitted.
Would love to hear more about that building energy modelling example, have you done a writeup you could share?
The Mamba application is my current research project so I haven’t published anything yet. But the basic idea is to create a latent representation of the static features, repeat the latent vector to form a time series, concatenate with the weather/occupancy time series, run through mamba layers, and bob’s your uncle. Shoot me an email (in my bio) if you would like to chat more! I can also share my master’s thesis whic…
I'm currently learning about machine learning and digital twin but don't really where to start
Re: A pure NumPy implementation of Mamba
#19Contrary to what the title says, this is note a pure Pyhon + numpy implementation: in fact it also imports einops, transformers and torch. For me pure X means: to use this, all you have to install is X.
"Yes, the comment you mentioned is fair and reflects a common perspective in the programming and data science communities regarding the usage of "pure" implementations. When someone refers to a "pure X implementation," the typical expectation is that the implementation will rely solely on the functionalities of library X, without introducing dependencies from other libraries or frameworks."
TIL.
Re: A pure NumPy implementation of Mamba
#20Contrary to what the title says, this is note a pure Pyhon + numpy implementation: in fact it also imports einops, transformers and torch. For me pure X means: to use this, all you have to install is X.
So it’s just numpy and einops, which is pretty cool. I guess you could probably rewrite all the einops stuff in pure numpy if you want to trade readable code for eliminating the einops dependency
Edit: found the torch import, but it’s just for a single torch.load to deserialize some data