Live data from Hacker News

Native Sparse Attention

aclanthology.org

11–20 of 32 posts

Re: Native Sparse Attention

#11
post #8

Deep seek papers are a must to read for anyone who wants to understand how to make LLMs operate at hyper scale. All western labs hide their best results, or at most release summaries that are about as meaningful as the answers Cleo used to give on stack exchange: https://math.stackexchange.com/questions/562694/integral-int... I have a suspicion with how quiet all the major players got after the two weeks after deepse…

None of the major players have ever been quiet. DeepSeek enjoyed about a week or two's worth of press before its spotlight was stolent from the next great model. It never held the top spot, ever, mind you. So I don't understand why you think major players had to say anything about it, when the model was neither first, second or third in real world capability, and why they would have to say anything about it when DeepSeek service processes maybe an 1/8 of what OpenAI, Google or Claude in any given span of time.

I applaud their open efforts. But being "altruistic" and being best are two different things.

Re: Native Sparse Attention

#12
post #11
post #8

Deep seek papers are a must to read for anyone who wants to understand how to make LLMs operate at hyper scale. All western labs hide their best results, or at most release summaries that are about as meaningful as the answers Cleo used to give on stack exchange: https://math.stackexchange.com/questions/562694/integral-int... I have a suspicion with how quiet all the major players got after the two weeks after deepse…

None of the major players have ever been quiet. DeepSeek enjoyed about a week or two's worth of press before its spotlight was stolent from the next great model. It never held the top spot, ever, mind you. So I don't understand why you think major players had to say anything about it, when the model was neither first, second or third in real world capability, and why they would have to say anything about it when Deep…

DeepSeek's contributions to training efficiency improvements were as, if not more, important than the models themselves. A lot of the worry people had about DeepSeek was related to people questioning the moat of the big AI players, since DeepSeek was able to train a competitive model with so much less compute.

Their innovations in training efficiency were almost guaranteed to have been heavily considered by the big AI labs. For example, Dario Amodei talks about the efficiency improvements being the real important contribution of DeepSeek V3 here: https://www.darioamodei.com/post/on-deepseek-and-export-cont...

> DeepSeek's team did this via some genuine and impressive innovations, mostly focused on engineering efficiency. There were particularly innovative improvements in the management of an aspect called the "Key-Value cache", and in enabling a method called "mixture of experts" to be pushed further than it had before.

Re: Native Sparse Attention

#13
I am always skeptical of RNN approaches but this paper is just sparsifying the input, it is not compressing any size input to a fixed memory. I am hopeful maybe this is a big break. 11x inference speedup with no degradation from an algorithmic improvement. Is it really that good? almost too good to be true. Adoption in the next 6 months will tell us the truth.

Re: Native Sparse Attention

#14
post #11
post #8

Deep seek papers are a must to read for anyone who wants to understand how to make LLMs operate at hyper scale. All western labs hide their best results, or at most release summaries that are about as meaningful as the answers Cleo used to give on stack exchange: https://math.stackexchange.com/questions/562694/integral-int... I have a suspicion with how quiet all the major players got after the two weeks after deepse…

None of the major players have ever been quiet. DeepSeek enjoyed about a week or two's worth of press before its spotlight was stolent from the next great model. It never held the top spot, ever, mind you. So I don't understand why you think major players had to say anything about it, when the model was neither first, second or third in real world capability, and why they would have to say anything about it when Deep…

MLA is just one example of a best-in-class technique from Hangzhou that's seen wide adoption in US prestige labs.

And the saltiness of US labs about DeepSeek is well-known. "O3, explain model distillation like I'm five."

No Sam, explain intellectual property rights to the judge in the NYT test case asshole.

Re: Native Sparse Attention

#15
post #8

Deep seek papers are a must to read for anyone who wants to understand how to make LLMs operate at hyper scale. All western labs hide their best results, or at most release summaries that are about as meaningful as the answers Cleo used to give on stack exchange: https://math.stackexchange.com/questions/562694/integral-int... I have a suspicion with how quiet all the major players got after the two weeks after deepse…

I remember on february Deepseek's caused a moderately sized market crash. They didn't just go silent, almost every vendor implemented their own version of thinking models while blaming Deepseek for stealing their tech/training on their models. It was rather pathetic to watch.

Re: Native Sparse Attention

#16

> Despite being sparse, NSA surpasses Full Attention baseline on average across general benchmarks, long-context tasks, and reasoning evaluation. Isn't it very notable that the latency improvement didn't have a performance loss? I'm not super familiar with all the technical aspects, but that seems like it should be one of the main focuses of the paper.

The performance maintenance (or even improvement) isn't surprising - sparse attention can reduce noise by focusing only on relevant tokens. Traditional full attention dilutes focus by attending to everything equally, while NSA's pruning approach mimics how humans selectively process information.

Re: Native Sparse Attention

#18
post #11
post #8

Deep seek papers are a must to read for anyone who wants to understand how to make LLMs operate at hyper scale. All western labs hide their best results, or at most release summaries that are about as meaningful as the answers Cleo used to give on stack exchange: https://math.stackexchange.com/questions/562694/integral-int... I have a suspicion with how quiet all the major players got after the two weeks after deepse…

None of the major players have ever been quiet. DeepSeek enjoyed about a week or two's worth of press before its spotlight was stolent from the next great model. It never held the top spot, ever, mind you. So I don't understand why you think major players had to say anything about it, when the model was neither first, second or third in real world capability, and why they would have to say anything about it when Deep…

Genuinely many times it seems most people need to find reasons to assume the best about DeepSeek and China in order to confirm their prior bias that “America bad” and “Capital is evil”. The reality is grey and fuzzy, with neither side landing on truth yet

Re: Native Sparse Attention

#19
post #11

Earlier quoted context omitted.

None of the major players have ever been quiet. DeepSeek enjoyed about a week or two's worth of press before its spotlight was stolent from the next great model. It never held the top spot, ever, mind you. So I don't understand why you think major players had to say anything about it, when the model was neither first, second or third in real world capability, and why they would have to say anything about it when Deep…

DeepSeek's contributions to training efficiency improvements were as, if not more, important than the models themselves. A lot of the worry people had about DeepSeek was related to people questioning the moat of the big AI players, since DeepSeek was able to train a competitive model with so much less compute. Their innovations in training efficiency were almost guaranteed to have been heavily considered by the big A…

Almost all of High Flyers achievements have more to do with scaling the process but when scaling is all you need, it’s darn effective

Re: Native Sparse Attention

#20
post #11

Earlier quoted context omitted.

None of the major players have ever been quiet. DeepSeek enjoyed about a week or two's worth of press before its spotlight was stolent from the next great model. It never held the top spot, ever, mind you. So I don't understand why you think major players had to say anything about it, when the model was neither first, second or third in real world capability, and why they would have to say anything about it when Deep…

MLA is just one example of a best-in-class technique from Hangzhou that's seen wide adoption in US prestige labs. And the saltiness of US labs about DeepSeek is well-known. "O3, explain model distillation like I'm five." No Sam, explain intellectual property rights to the judge in the NYT test case asshole.

… wait did you just seriously tell SamA that he’s an asshole because of copyright issues… while praising Chinese labs who couldn’t give a rat fuck and won’t follow the same laws? Or pay creators? Physician, heal thyself
Post reply on HN