Live data from Hacker News

Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model

twitter.com

51–60 of 194 posts

Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model

#51

Earlier quoted context omitted.

Inseparable, routing is done per token in a statistically optimal way, not per request on the knowledge domain basis.

Sure, it's done per token, but the question is: how much do the knowledge domains match up with experts. I could not find hard data on this.

Check out DeepSeek v3 model paper. They changed the way they train experts (went from aux loss to different kind expert separation training). It did improve experts domain specialization, they have neat graphics on it in the paper.

Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model

#52
post #39

Earlier quoted context omitted.

I'm glad we are looking to build nuclear reactors so we can do more of this...

me too - we must energymaxx. i want a nuclear reactor in my backyard powering everything. I want ac units in every room and my open door garage while i workout.

You're saying this in jest, but I would LOVE to have a nuclear reactor in my backyard that produced enough power to where I could have a minisplit for every room in my house, including the garage so I could work out in there.

Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model

#54
post #29
post #18

Would be hilarious if Zuck with his billion dollar poaching failed to beat budget Chinese models.

That reminds me of a thought I had about the poachings. The poaching was probably more aimed at hamstringing Meta's competition. Because the disruption caused by them leaving in droves is probably more severe than the benefits of having them on board. Unless they are gods, of course.

I thought that too

Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model

#55

Earlier quoted context omitted.

me too - we must energymaxx. i want a nuclear reactor in my backyard powering everything. I want ac units in every room and my open door garage while i workout.

You're saying this in jest, but I would LOVE to have a nuclear reactor in my backyard that produced enough power to where I could have a minisplit for every room in my house, including the garage so I could work out in there.

Related: https://en.wikipedia.org/wiki/Kardashev_scale

> The Kardashev scale (Russian: шкала Кардашёва, romanized: shkala Kardashyova) is a method of measuring a civilization's level of technological advancement based on the amount of energy it is capable of harnessing and using.

> Under this scale, the sum of human civilization does not reach Type I status, though it continues to approach it.

Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model

#57
This is a very impressive general purpose LLM (GPT 4o, DeepSeek-V3 family). It’s also open source.

I think it hasn’t received much attention because the frontier shifted to reasoning and multi-modal AI models. In accuracy benchmarks, all the top models are reasoning ones:

https://artificialanalysis.ai/

If someone took Kimi k2 and trained a reasoning model with it, I’d be curious how that model performs.

Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model

#58
post #57

This is a very impressive general purpose LLM (GPT 4o, DeepSeek-V3 family). It’s also open source. I think it hasn’t received much attention because the frontier shifted to reasoning and multi-modal AI models. In accuracy benchmarks, all the top models are reasoning ones: https://artificialanalysis.ai/ If someone took Kimi k2 and trained a reasoning model with it, I’d be curious how that model performs.

>If someone took Kimi k2 and trained a reasoning model with it

I imagine that's what they are going at MoonshotAI right now

Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model

#59
This is the model release that made Sam Altman go "Oh wait actually we can't release the new open source model this week, sorry. Something something security concerns".

Perhaps their open source model release doesn't look so good compared to this one

Post reply on HN