Flash-MSA: Accelerating Million-Token Training with Sparse Attention Kernels
nanduruganesh.github.io
Flash-MSA: Accelerating Million-Token Training with Sparse Attention Kernels
1–6 of 6 posts
Re: Flash-MSA: Accelerating Million-Token Training with Sparse Attention Kernels
#2Such lazy, much farming
https://github.com/fla-org/native-sparse-attention?utm_sourc...
Re: Flash-MSA: Accelerating Million-Token Training with Sparse Attention Kernels
#3World’s first? Such lazy, much farming https://github.com/fla-org/native-sparse-attention?utm_sourc...
Re: Flash-MSA: Accelerating Million-Token Training with Sparse Attention Kernels
#4Has anyone used the new Minimax M3 model? I’m curious how it compares with Deepseek V4 and GLM 5.2 and other larger open weights models.
Re: Flash-MSA: Accelerating Million-Token Training with Sparse Attention Kernels
#5I’ve actually been really interested in Minimax M3 - seems like it flew under the radar but size wise might actually be runnable for local inference with a footprint somewhere between Deepseek V4 flash and pro. Has anyone used the new Minimax M3 model? I’m curious how it compares with Deepseek V4 and GLM 5.2 and other larger open weights models.
[Update: their cheapest token plan has been removed, I guess its back to GLM now]
Right now M3 is not far behind DS4, but I belive DS4 will improve much more with each round of training. It simply has a bigger brain, it just needs to fill it with more information.