Live data from Hacker News

A walk through of the DeltaNet family of linear attention variants

blog.doubleword.ai

1–10 of 134 posts

Re: A walk through of the DeltaNet family of linear attention variants

#8
I could never get this about modern machine/deep learning or even the Transformers. Yes, it's not exactly rocket science, but when I see the data flow diagrams, it's not clear what is calculated in real time or multiple steps.

Is it really one big computation f(g(h(x)))?

Post reply on HN