Live data from Hacker News

Native Sparse Attention

aclanthology.org

21–30 of 32 posts

Re: Native Sparse Attention

#21
post #8

Deep seek papers are a must to read for anyone who wants to understand how to make LLMs operate at hyper scale. All western labs hide their best results, or at most release summaries that are about as meaningful as the answers Cleo used to give on stack exchange: https://math.stackexchange.com/questions/562694/integral-int... I have a suspicion with how quiet all the major players got after the two weeks after deepse…

I remember on february Deepseek's caused a moderately sized market crash. They didn't just go silent, almost every vendor implemented their own version of thinking models while blaming Deepseek for stealing their tech/training on their models. It was rather pathetic to watch.

OAI and others were already on their way there or released the models. How did you manage to convince yourself that High Flyer did it first ? And that everyone else copied from them post-hoc? You’ve created a new chain of causality that simply does not match neutral reality

Re: Native Sparse Attention

#22

> Despite being sparse, NSA surpasses Full Attention baseline on average across general benchmarks, long-context tasks, and reasoning evaluation. Isn't it very notable that the latency improvement didn't have a performance loss? I'm not super familiar with all the technical aspects, but that seems like it should be one of the main focuses of the paper.

Yes that’s what makes it so interesting and novel you nailed it

Re: Native Sparse Attention

#24

Earlier quoted context omitted.

I remember on february Deepseek's caused a moderately sized market crash. They didn't just go silent, almost every vendor implemented their own version of thinking models while blaming Deepseek for stealing their tech/training on their models. It was rather pathetic to watch.

OAI and others were already on their way there or released the models. How did you manage to convince yourself that High Flyer did it first ? And that everyone else copied from them post-hoc? You’ve created a new chain of causality that simply does not match neutral reality

Yeah I confess I rewrote history and crashed the stock market. Then ran out of juice just as I was about to kill Hitler.

Re: Native Sparse Attention

#25

Earlier quoted context omitted.

MLA is just one example of a best-in-class technique from Hangzhou that's seen wide adoption in US prestige labs. And the saltiness of US labs about DeepSeek is well-known. "O3, explain model distillation like I'm five." No Sam, explain intellectual property rights to the judge in the NYT test case asshole.

… wait did you just seriously tell SamA that he’s an asshole because of copyright issues… while praising Chinese labs who couldn’t give a rat fuck and won’t follow the same laws? Or pay creators? Physician, heal thyself

Sam's an asshole for a lot of reasons, a ridiculous commons grab of intellectual property draped in threadbare rhetoric about human welfare (get those developing nation eyeballs SCANNED people!) being just one of them.

Watching the Chinese labs kick the shit out of better funded US enclaves of TESCREAL psychopathy in the public fucking domain is gravy.

I don't care that their internal calculus or that of the PRC is to Cloud Strife Limit Break a bunch of "shareholder value" in the form of a bloated NVIDIA cap feeding frenzy by bloated "public benefit corporations" with a bunch of creepy ties to Thiel et al: they're publishing papers, code and weights. So they're hoovering up of the commons has something of value going back into the commons.

So yeah, fuck Sam and its going to be fun watching OpenAI and Anthropic pivot ever more towards trying to outlaw competition than they already have. Amodei already sounds like Donald Rumsfeld on Taiwan hawkishness, this is not the positioning of someone who loves their product roadmap.

It turns out that a zillion ScaleAI and SurgeAI turks don't have economics any better than paying NVIDIA to run 85% net earnings for CapEx that's obsolete by the time its racked and powered.

Re: Native Sparse Attention

#26
post #11

Earlier quoted context omitted.

None of the major players have ever been quiet. DeepSeek enjoyed about a week or two's worth of press before its spotlight was stolent from the next great model. It never held the top spot, ever, mind you. So I don't understand why you think major players had to say anything about it, when the model was neither first, second or third in real world capability, and why they would have to say anything about it when Deep…

Genuinely many times it seems most people need to find reasons to assume the best about DeepSeek and China in order to confirm their prior bias that “America bad” and “Capital is evil”. The reality is grey and fuzzy, with neither side landing on truth yet

How would people use deepseek to think "Capital is evil?" It was from a private hedge fund named "High Flyer," not a state university project or something.

Re: Native Sparse Attention

#27

Earlier quoted context omitted.

OAI and others were already on their way there or released the models. How did you manage to convince yourself that High Flyer did it first ? And that everyone else copied from them post-hoc? You’ve created a new chain of causality that simply does not match neutral reality

Yeah I confess I rewrote history and crashed the stock market. Then ran out of juice just as I was about to kill Hitler.

Do not try to signal intelligence by being sardonic or intentionally being obtuse. Actively avoiding the point someone is making rather than confronting it head on is beneath you.

ChatGPT o1 was made generally available in December 2024 DeepSeek r1 open weights were released in January 2025

Re: Native Sparse Attention

#28

Earlier quoted context omitted.

… wait did you just seriously tell SamA that he’s an asshole because of copyright issues… while praising Chinese labs who couldn’t give a rat fuck and won’t follow the same laws? Or pay creators? Physician, heal thyself

Sam's an asshole for a lot of reasons, a ridiculous commons grab of intellectual property draped in threadbare rhetoric about human welfare (get those developing nation eyeballs SCANNED people!) being just one of them. Watching the Chinese labs kick the shit out of better funded US enclaves of TESCREAL psychopathy in the public fucking domain is gravy. I don't care that their internal calculus or that of the PRC is t…

... You did not speak to the key point at all and went on some massive rambling incoherent political commentary. I feel this comment is unworthy of the thread.

Native Sparse Attention matters. Your commentary is beneath this paper.

Re: Native Sparse Attention

#29
post #26

Earlier quoted context omitted.

Genuinely many times it seems most people need to find reasons to assume the best about DeepSeek and China in order to confirm their prior bias that “America bad” and “Capital is evil”. The reality is grey and fuzzy, with neither side landing on truth yet

How would people use deepseek to think "Capital is evil?" It was from a private hedge fund named "High Flyer," not a state university project or something.

Yes, exactly. How the heck? It makes no sense to me either, but you can certainly find plenty of laymen/not-in-the-know folks making those kinds of comments, often in non-technical spaces. Often the worst parts of the internet where discourse is non-existent. Human psychology allows for us to hold many contradictory positions all at once. Ideologies are the lens through which we view the world and it distorts our perception.

Re: Native Sparse Attention

#30
post #11
post #8

Deep seek papers are a must to read for anyone who wants to understand how to make LLMs operate at hyper scale. All western labs hide their best results, or at most release summaries that are about as meaningful as the answers Cleo used to give on stack exchange: https://math.stackexchange.com/questions/562694/integral-int... I have a suspicion with how quiet all the major players got after the two weeks after deepse…

None of the major players have ever been quiet. DeepSeek enjoyed about a week or two's worth of press before its spotlight was stolent from the next great model. It never held the top spot, ever, mind you. So I don't understand why you think major players had to say anything about it, when the model was neither first, second or third in real world capability, and why they would have to say anything about it when Deep…

It crashed the market because retail investors and perhaps non-retail as well had a great deal in overconfidence with the ability of the USA to maintain a lead thanks to the chip gap. High Flyer's innovations allowed them to scale and show that is not the case. This major event then likely spurred on many others. It was a mini 'sputnik moment'
Post reply on HN