A walk through of the DeltaNet family of linear attention variants
51–60 of 134 posts
Re: A walk through of the DeltaNet family of linear attention variants
#52why are you using braket notation?
He has a master's degree in physics from Oxford. Also there is a toggle to normal notation. Well, CS notation. I'm not a fan of transpose marks everywhere. I like an even more mathematics notation.
Re: A walk through of the DeltaNet family of linear attention variants
#53after a cursory read, I can confidently say I could not, in fact, have come up with Kimi Delta Attention.
Re: A walk through of the DeltaNet family of linear attention variants
#54after a cursory read, I can confidently say I could not, in fact, have come up with Kimi Delta Attention.
"you could have..." is among the top insulting phrases used by maths-adjacent people. Others in that league are "it should now be obvious...", "it's abundantly clear...", "it can be easily shown that...", "this is nothing but..." etc. The rest of us reading this are like, holy batman, what the fuck was that?!
Re: A walk through of the DeltaNet family of linear attention variants
#55Re: A walk through of the DeltaNet family of linear attention variants
#56after a cursory read, I can confidently say I could not, in fact, have come up with Kimi Delta Attention.
"You could have come up with Kimi Delta Attention, but you didn't, did you."
Re: A walk through of the DeltaNet family of linear attention variants
#57I would have liked some refresher on some variables though (like d_k in quadratic attention).
Re: A walk through of the DeltaNet family of linear attention variants
#58Machine learning could need, and probably has needed, some unified math notation for the past 15 years IMO. With that said, it was worse back in the day - when ML papers were the products of researchers from all over, you'd see some wild notation. Many will likely disagree with me, but inconsistent notation (across papers!) is to me friction. At least in this article the author explicitly explains the notation at the…
Re: A walk through of the DeltaNet family of linear attention variants
#59Re: A walk through of the DeltaNet family of linear attention variants
#60Machine learning could need, and probably has needed, some unified math notation for the past 15 years IMO. With that said, it was worse back in the day - when ML papers were the products of researchers from all over, you'd see some wild notation. Many will likely disagree with me, but inconsistent notation (across papers!) is to me friction. At least in this article the author explicitly explains the notation at the…
I never understood people who preferred traditional math notation (e.g. single letter symbols, weird characters like ∣q⟩ instead of writing down an explicit type, etc.). I guess the main advantage is terseness? To me, the mathematical expressions would be so much easier to understand if they were just written in pseudo code or an actual programming language like Python.