Live data from Hacker News

A walk through of the DeltaNet family of linear attention variants

blog.doubleword.ai

131–134 of 134 posts

Re: A walk through of the DeltaNet family of linear attention variants

#131
post #60

Earlier quoted context omitted.

I never understood people who preferred traditional math notation (e.g. single letter symbols, weird characters like ∣q⟩ instead of writing down an explicit type, etc.). I guess the main advantage is terseness? To me, the mathematical expressions would be so much easier to understand if they were just written in pseudo code or an actual programming language like Python.

math notation doesn't bias towards English language understanding like pseudocode

Diversity win: this notation can't be easily understood by anyone!

Re: A walk through of the DeltaNet family of linear attention variants

#132
post #85
post #75

Earlier quoted context omitted.

Doubleword AI is conducting a classic textbook marketing trick called newsjacking. Writing a detailed technical post behind the news of Kimi K3 and KDA algorithm with an audacious title like "You Could Have Invent Breakthrough It too" they are pre-filtering out the ones who couldn't comprehend with quick read (myself included) and attracting the ones who agreed with the blog post. At the end with a strong CTA to prom…

“Kimi Delta Attention” because “Kimi K3 Delta Attention (oh that’s just our little internal name for it as a joke)” passes no sniff tests.

Nevermind, too thick to understand this

Re: A walk through of the DeltaNet family of linear attention variants

#133
post #121
post #101

Earlier quoted context omitted.

Well, the goal was to make them come up with cutting-edge theories and/or hypothesis to the most pressing problems facing humanity. Kids will naturally specialize (just like... MoE models?) as their find their groove while growing up, but the goal is to make sure highly gifted kids from diverse background get highest quality space to think without being moulded or "adulterated" by the noise of pop culture and "everyd…

> make sure highly gifted kids from diverse background It would be ideal to get those gifted kids before they spend 5-6 years in the 'real world' with all the "Baby Shark," Roblox, and various other brainrot, but if you pick kids randomly from a very early age, most of them will probably not be all that gifted. Though I suppose maybe you could rely on genetic heritage (parents of very high IQ) and/or just do a large…

They’re going to be so mad when they graduate The Program and find out about Minecraft

Re: A walk through of the DeltaNet family of linear attention variants

#134
post #38

Earlier quoted context omitted.

bra-ket is the (most?) general form of tensor manipulation. Raising and lowering operators for summation notation are the beginner tools for covariant derivatives of the metric tensor. Christoffel symbols are where it's at, if you need to write out the Ricci tensor. The more constrained the space the more concise the notation can be. Note that MechE tensor notation has an even more compact (eigen) form for principal…

I don't think there is anything in this article that actually demands bra-ket notation (a state in some Hilbert space), that couldn't be more clearly written with standard notation for a Euclidean inner/outer product, but I suppose everyone has their own preferences for notation.

The article even allows you to toggle the notation used.
Post reply on HN