Live data from Hacker News

Native Sparse Attention

aclanthology.org

1–10 of 32 posts

Re: Native Sparse Attention

#3
> Despite being sparse, NSA surpasses Full Attention baseline on average across general benchmarks, long-context tasks, and reasoning evaluation.

Isn't it very notable that the latency improvement didn't have a performance loss? I'm not super familiar with all the technical aspects, but that seems like it should be one of the main focuses of the paper.

Re: Native Sparse Attention

#5
post #4

Title: Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention The awards page for ACL seems to disagree with this editorialized title: https://2025.aclweb.org/program/awards/

The ACL webpage has not been updated yet. Here are the announcement slides: https://cspaper.org/topic/116/record-breaking-acl-2025-crown...

Re: Native Sparse Attention

#6
I'd say award for best title is a tie between: "Dehumanizing Machines: Mitigating Anthropomorphic Behaviors in Text Generation Systems"; "Finding Needles in Images: Can Multi-modal LLMs Locate Fine Details?"; and "Steering off Course: Reliability Challenges in Steering Language Models."

Re: Native Sparse Attention

#8
Deep seek papers are a must to read for anyone who wants to understand how to make LLMs operate at hyper scale. All western labs hide their best results, or at most release summaries that are about as meaningful as the answers Cleo used to give on stack exchange: https://math.stackexchange.com/questions/562694/integral-int...

I have a suspicion with how quiet all the major players got after the two weeks after deepseek R1 was released that they were reading and implementing everything in the papers that came with it as fast as humanly possible.

Re: Native Sparse Attention

#10
post #5
post #4

Title: Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention The awards page for ACL seems to disagree with this editorialized title: https://2025.aclweb.org/program/awards/

The ACL webpage has not been updated yet. Here are the announcement slides: https://cspaper.org/topic/116/record-breaking-acl-2025-crown...

The page that the person you’re replying to does have this so it may not be updated, or they were looking in the wrong place originally, or both:

> Industry Track Awards

> Best Paper

> Speed Without Sacrifice: Fine-Tuning Language Models with Medusa and Knowledge Distillation in Travel Applications

> Daniel Zagyva, Emmanouil Stergiadis, Laurens van der Maas, Aleksandra Dokic, Eran Fainman, Ilya Gusev, Moran Beladev

Per TFA, the paper we’re looking for is this one:

> Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

> Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Zhengyan Zhang, Zhenda Xie, Y. X. Wei, Lean Wang, Zhiping Xiao, Yuqing Wang, Chong Ruan, Ming Zhang, Wenfeng Liang, Wangding Zeng

I’m not finding it by author on the page you linked but I think it’s this reference by title:

> DeepSeek × PKU × UW — Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

I did find it on this page:

https://2025.aclweb.org/program/main_papers/

Post reply on HN