If you want to believe that the success of Kimi is about distillation attacks, ignore this.
Kimi Linear: An Expressive, Efficient Attention Architecture (2025)
21–30 of 141 posts
Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)
#22This is just awesome.
Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)
#23Does any expert in the field know whether it is really the case that this intelligence we are seeing with frontier models is an "emerging" phenomena, only coming up when the architecture is scaled? Like isn't it weird that the 1 million parameter model with the same architecture can't solve basic puzzles but suddenly the 1 trillion parameter can conjure up counter-examples for the Jacobian conjecture? It's unintuitiv…
Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)
#24Does any expert in the field know whether it is really the case that this intelligence we are seeing with frontier models is an "emerging" phenomena, only coming up when the architecture is scaled? Like isn't it weird that the 1 million parameter model with the same architecture can't solve basic puzzles but suddenly the 1 trillion parameter can conjure up counter-examples for the Jacobian conjecture? It's unintuitiv…
Let's say there's some circuit that does problem solving of the kind we call intelligence.
We dont know what this circuit looks like, but it exists in our brain.
Doing regression on outputs from the brain (e.g. internet text) with enough parameters, we can "fit" our model to this circuit.
But if you try to fit it with fewer parameters than it needs, you're just going to get some linear approximation.
Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)
#25Does any expert in the field know whether it is really the case that this intelligence we are seeing with frontier models is an "emerging" phenomena, only coming up when the architecture is scaled? Like isn't it weird that the 1 million parameter model with the same architecture can't solve basic puzzles but suddenly the 1 trillion parameter can conjure up counter-examples for the Jacobian conjecture? It's unintuitiv…
https://en.wikipedia.org/wiki/Grokking_(machine_learning)
High-dimensional gradient descent behaves very differently than the simplified 3d visualisations we use to demonstrate it, and has lots of ways out of local minima:
https://www.youtube.com/watch?v=NrO20Jb-hy0
so it seems like there is a benefit to giving models more space to learn in rather than forcing them to compress the knowledge from the start.
Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)
#26Does any expert in the field know whether it is really the case that this intelligence we are seeing with frontier models is an "emerging" phenomena, only coming up when the architecture is scaled? Like isn't it weird that the 1 million parameter model with the same architecture can't solve basic puzzles but suddenly the 1 trillion parameter can conjure up counter-examples for the Jacobian conjecture? It's unintuitiv…
Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)
#27Does any expert in the field know whether it is really the case that this intelligence we are seeing with frontier models is an "emerging" phenomena, only coming up when the architecture is scaled? Like isn't it weird that the 1 million parameter model with the same architecture can't solve basic puzzles but suddenly the 1 trillion parameter can conjure up counter-examples for the Jacobian conjecture? It's unintuitiv…
> One thing that should be learned from the bitter lesson is the great power of general purpose methods, of methods that continue to scale with increased computation even as the available computation becomes very great. The two methods that seem to scale arbitrarily in this way are search and learning.
The full essay is worth a read, it's pretty short http://www.incompleteideas.net/IncIdeas/BitterLesson.html
Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)
#28If you want to believe that the success of Kimi is about distillation attacks, ignore this.
False dichotomy right? Are Chinese labs impressively innovating? Clearly. However this doesn’t rule out possible gains due to distillation. I don’t know the degree of the latter but both things could certainly be true.
“You stole my warez!”
Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)
#29Does any expert in the field know whether it is really the case that this intelligence we are seeing with frontier models is an "emerging" phenomena, only coming up when the architecture is scaled? Like isn't it weird that the 1 million parameter model with the same architecture can't solve basic puzzles but suddenly the 1 trillion parameter can conjure up counter-examples for the Jacobian conjecture? It's unintuitiv…
Going from 1m to 1T params is also a scaling of a million times. It’s like going from a human brain down to one percent in size in each direction or just a few mm.
Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)
#30If you want to believe that the success of Kimi is about distillation attacks, ignore this.
The distillation theory does not even make sense as Fable was only around for days (effectively) before Kimi was released.