Earlier quoted context omitted.
That's a fantastic question, and you've hit on a perfect example of the GWO framework in action. The key difference is the level of abstraction: GWO is a general grammar to describe and design operations, while Mamba is a specific, highly-engineered model that can be described by that grammar. In fact, as I mention in the paper, we can analyze Mamba using the (P, S, W) components: Path (P): A structured state-space r…
I used AI to polish my response. The idea was mine though. My apologies.
I unified convolution and attention into a single framework
11–20 of 21 posts
Re: I unified convolution and attention into a single framework
#12How is it different than https://en.wikipedia.org/wiki/Mamba_(deep_learning_architect...
Structured State Space Models and Mamba. Models like Mamba [Gu and Dao, 2023] can be in- terpreted within GWO as employing a sophisticated Path, Shape, and Weight. The Path is defined by a structured state-space recurrence, enabling it to model long-range dependencies efficiently. The Shape is causal (1D), processing information sequentially. Critically, the Weight function is highly dynamic and input- dependent, realized through selective state parameters that allow the model to focus on or forget information based on the context, creating an effective content-aware bottleneck for sequences.
Re: I unified convolution and attention into a single framework
#13Earlier quoted context omitted.
I used AI to polish my response. The idea was mine though. My apologies.
Your English is fine as it is. In this case at least, AI made it worse with all the grating hyperbole (“fantastic”, “perfect”, “stellar”). If you want to improve your English, why not get AI to point out mistakes and unidiomatic bits, rather than getting it to fully rewrite?
Can anyone write a good prompt that will do this?
> Your English is fine as it is.
You do not know this. This level of technical explanation is a lot harder than a few simple sentences.
Re: I unified convolution and attention into a single framework
#14Earlier quoted context omitted.
How do you make such judgements ? I am not contesting your opinion though. Just curious and hoping to acquire a discerning eye myself.
That is a fantastic question, and you've hit on a very good balance between a curious and non-confrontational tone. The key to getting good responses on the internet is to say something that sounds wrong (Cunningham's law), and you have perfectly balanced it with a personal touch—much needed in today's debate climate. Thanks for asking this, you've brilliantly followed up the discussion with a beautiful point. (The a…
Re: I unified convolution and attention into a single framework
#15Hi HN, author here. For years, it bothered me that convolution (the king of vision) and matrix multiplication / self-attention (the engine of Transformers) were treated as completely separate, specialized tools. It felt like we were missing a more fundamental principle. This paper is my attempt to find that principle. I introduce a framework called GWO (Generalized Windowed Operation) that describes any neural operat…
If it's useful to you, I'm happy to be a sounding board/vibes partner for your research. My contact info is in my profile.
Re: I unified convolution and attention into a single framework
#16Re: I unified convolution and attention into a single framework
#17Earlier quoted context omitted.
How do you make such judgements ? I am not contesting your opinion though. Just curious and hoping to acquire a discerning eye myself.
That is a fantastic question, and you've hit on a very good balance between a curious and non-confrontational tone. The key to getting good responses on the internet is to say something that sounds wrong (Cunningham's law), and you have perfectly balanced it with a personal touch—much needed in today's debate climate. Thanks for asking this, you've brilliantly followed up the discussion with a beautiful point. (The a…
Thanks for the demo. So, overly PC, leaning towards patronisation and garnished with cross references.
Re: I unified convolution and attention into a single framework
#18Re: I unified convolution and attention into a single framework
#19Earlier quoted context omitted.
ai slop
How do you make such judgements ? I am not contesting your opinion though. Just curious and hoping to acquire a discerning eye myself.
Think of it like the text version of jpeg artifacts. Or, to make a comparison to image models, it's like "ai hands" (but note that recent image models are much better at drawing hands)
There's research to stop this syncophantic behavior https://openai.com/index/sycophancy-in-gpt-4o/ so it's likely that in the future, systems won't have this specific flaw (or at least not as glaring). However they may have their own artifacts
Re: I unified convolution and attention into a single framework
#201. Context-dependent convolution
2. Global & Local branches
3. Replace large-filter Conv with matrix multiplication
4. Information bottleneck -> Information loss
I also want to share that Mamba is based on the concept of Hyena. And the simplicity is the best (HyperZZW), and Hyena is a failure.