"Interestingly, Kimi K3 got rid of all RoPE layers and uses NoPE (No Positional Embeddings) everywhere instead." It just baffles me that this even works at all. Doesn't it just become a token soup? Is attention that precise that a second token can tell its the second token just because it learns to accumulate something in the embedding space without any sort of inductive bias?
Kimi K3 Architecture Overview and Notes
51–60 of 125 posts
Re: Kimi K3 Architecture Overview and Notes
#52"Interestingly, Kimi K3 got rid of all RoPE layers and uses NoPE (No Positional Embeddings) everywhere instead." It just baffles me that this even works at all. Doesn't it just become a token soup? Is attention that precise that a second token can tell its the second token just because it learns to accumulate something in the embedding space without any sort of inductive bias?
Re: Kimi K3 Architecture Overview and Notes
#53My only doubts are around Linear Attention instead of DSA as this is inherently lossy. You are kind of banking on that your query is inherently in the embedding space of the model already and can be lossy.
Re: Kimi K3 Architecture Overview and Notes
#54So, unlike what leaders of western labs labs would like you to believe (that Kimi is just the result of distillation attacks), they are introducing new and novel approaches.
"Kimi is largely a byproduct of distillation" and "Kimi is introducing new and novel approaches" are not mutually exclusive, and I'm not sure it's clear from the paper how much of the improvement comes from the new approaches. So I wouldn't take the new approaches to be much evidence about whether the distillation attacks occurred.
Like the distillation “attacks” Anthropic conducted against literally millions and millions of different creators?
Re: Kimi K3 Architecture Overview and Notes
#55Earlier quoted context omitted.
Even if they are distilling, I don't particularly care. Anthropic and others have been distilling copyrighted material by to build these models, largely without permission.
[flagged]
Re: Kimi K3 Architecture Overview and Notes
#56Earlier quoted context omitted.
"Kimi is largely a byproduct of distillation" and "Kimi is introducing new and novel approaches" are not mutually exclusive, and I'm not sure it's clear from the paper how much of the improvement comes from the new approaches. So I wouldn't take the new approaches to be much evidence about whether the distillation attacks occurred.
> distillation attacks Like the distillation “attacks” Anthropic conducted against literally millions and millions of different creators?
Re: Kimi K3 Architecture Overview and Notes
#57I feel like the Kimi team is amongst the best in the industry to pick and choose what is meaningful from the other models. For example, avoiding the expensive and empirically uncertain mHC in favor of simpler residuals. Latent MoE. My only doubts are around Linear Attention instead of DSA as this is inherently lossy. You are kind of banking on that your query is inherently in the embedding space of the model already…
Re: Kimi K3 Architecture Overview and Notes
#58Earlier quoted context omitted.
Even if they are distilling, I don't particularly care. Anthropic and others have been distilling copyrighted material by to build these models, largely without permission.
[flagged]
Re: Kimi K3 Architecture Overview and Notes
#59Earlier quoted context omitted.
[flagged]
People don't remember now, but in the old HN comments like "this is so true honestly" would have been downvoted. If a comment just agrees with the parent, people would have said that's what the upvote button is for. Adding a "I agree" comment is just noise.
Re: Kimi K3 Architecture Overview and Notes
#60Earlier quoted context omitted.
People don't remember now, but in the old HN comments like "this is so true honestly" would have been downvoted. If a comment just agrees with the parent, people would have said that's what the upvote button is for. Adding a "I agree" comment is just noise.
This is so true honestly.