A walk through of the DeltaNet family of linear attention variants
blog.doubleword.ai
A walk through of the DeltaNet family of linear attention variants
1–10 of 134 posts
Re: A walk through of the DeltaNet family of linear attention variants
#2after a cursory read, I can confidently say I could not, in fact, have come up with Kimi Delta Attention.
Re: A walk through of the DeltaNet family of linear attention variants
#3after a cursory read, I can confidently say I could not, in fact, have come up with Kimi Delta Attention.
Not even close for me too.
Re: A walk through of the DeltaNet family of linear attention variants
#4[flagged]
Re: A walk through of the DeltaNet family of linear attention variants
#5after a cursory read, I can confidently say I could not, in fact, have come up with Kimi Delta Attention.
I thought I was the only one.
Re: A walk through of the DeltaNet family of linear attention variants
#6why are you using braket notation?
Re: A walk through of the DeltaNet family of linear attention variants
#7I like the math vs physics toggle.
Re: A walk through of the DeltaNet family of linear attention variants
#8I could never get this about modern machine/deep learning or even the Transformers. Yes, it's not exactly rocket science, but when I see the data flow diagrams, it's not clear what is calculated in real time or multiple steps.
Is it really one big computation f(g(h(x)))?
Re: A walk through of the DeltaNet family of linear attention variants
#9after a cursory read, I can confidently say I could not, in fact, have come up with Kimi Delta Attention.
I don't even know most words they used in the paper haha
Re: A walk through of the DeltaNet family of linear attention variants
#10You know its a doozy when the author writes a disclaimer at the top saying that bra-ket notation was chosen in order to make the algorithm and data structures clearer.