Earlier quoted context omitted.
Inseparable, routing is done per token in a statistically optimal way, not per request on the knowledge domain basis.
Sure, it's done per token, but the question is: how much do the knowledge domains match up with experts. I could not find hard data on this.
Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
51–60 of 194 posts
Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#52Earlier quoted context omitted.
I'm glad we are looking to build nuclear reactors so we can do more of this...
me too - we must energymaxx. i want a nuclear reactor in my backyard powering everything. I want ac units in every room and my open door garage while i workout.
Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#53Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#54Would be hilarious if Zuck with his billion dollar poaching failed to beat budget Chinese models.
That reminds me of a thought I had about the poachings. The poaching was probably more aimed at hamstringing Meta's competition. Because the disruption caused by them leaving in droves is probably more severe than the benefits of having them on board. Unless they are gods, of course.
Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#55Earlier quoted context omitted.
me too - we must energymaxx. i want a nuclear reactor in my backyard powering everything. I want ac units in every room and my open door garage while i workout.
You're saying this in jest, but I would LOVE to have a nuclear reactor in my backyard that produced enough power to where I could have a minisplit for every room in my house, including the garage so I could work out in there.
> The Kardashev scale (Russian: шкала Кардашёва, romanized: shkala Kardashyova) is a method of measuring a civilization's level of technological advancement based on the amount of energy it is capable of harnessing and using.
> Under this scale, the sum of human civilization does not reach Type I status, though it continues to approach it.
Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#56Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#57I think it hasn’t received much attention because the frontier shifted to reasoning and multi-modal AI models. In accuracy benchmarks, all the top models are reasoning ones:
https://artificialanalysis.ai/
If someone took Kimi k2 and trained a reasoning model with it, I’d be curious how that model performs.
Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#58This is a very impressive general purpose LLM (GPT 4o, DeepSeek-V3 family). It’s also open source. I think it hasn’t received much attention because the frontier shifted to reasoning and multi-modal AI models. In accuracy benchmarks, all the top models are reasoning ones: https://artificialanalysis.ai/ If someone took Kimi k2 and trained a reasoning model with it, I’d be curious how that model performs.
I imagine that's what they are going at MoonshotAI right now
Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#59Perhaps their open source model release doesn't look so good compared to this one
Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#60Open-weight. As usual, you don't get the dataset, training scripts, etc.