Earlier quoted context omitted.
I never understood people who preferred traditional math notation (e.g. single letter symbols, weird characters like ∣q⟩ instead of writing down an explicit type, etc.). I guess the main advantage is terseness? To me, the mathematical expressions would be so much easier to understand if they were just written in pseudo code or an actual programming language like Python.
math notation doesn't bias towards English language understanding like pseudocode
A walk through of the DeltaNet family of linear attention variants
131–134 of 134 posts
Re: A walk through of the DeltaNet family of linear attention variants
#132Earlier quoted context omitted.
Doubleword AI is conducting a classic textbook marketing trick called newsjacking. Writing a detailed technical post behind the news of Kimi K3 and KDA algorithm with an audacious title like "You Could Have Invent Breakthrough It too" they are pre-filtering out the ones who couldn't comprehend with quick read (myself included) and attracting the ones who agreed with the blog post. At the end with a strong CTA to prom…
“Kimi Delta Attention” because “Kimi K3 Delta Attention (oh that’s just our little internal name for it as a joke)” passes no sniff tests.
Re: A walk through of the DeltaNet family of linear attention variants
#133Earlier quoted context omitted.
Well, the goal was to make them come up with cutting-edge theories and/or hypothesis to the most pressing problems facing humanity. Kids will naturally specialize (just like... MoE models?) as their find their groove while growing up, but the goal is to make sure highly gifted kids from diverse background get highest quality space to think without being moulded or "adulterated" by the noise of pop culture and "everyd…
> make sure highly gifted kids from diverse background It would be ideal to get those gifted kids before they spend 5-6 years in the 'real world' with all the "Baby Shark," Roblox, and various other brainrot, but if you pick kids randomly from a very early age, most of them will probably not be all that gifted. Though I suppose maybe you could rely on genetic heritage (parents of very high IQ) and/or just do a large…
Re: A walk through of the DeltaNet family of linear attention variants
#134Earlier quoted context omitted.
bra-ket is the (most?) general form of tensor manipulation. Raising and lowering operators for summation notation are the beginner tools for covariant derivatives of the metric tensor. Christoffel symbols are where it's at, if you need to write out the Ricci tensor. The more constrained the space the more concise the notation can be. Note that MechE tensor notation has an even more compact (eigen) form for principal…
I don't think there is anything in this article that actually demands bra-ket notation (a state in some Hilbert space), that couldn't be more clearly written with standard notation for a Euclidean inner/outer product, but I suppose everyone has their own preferences for notation.