Pelican on a bicycle result: https://simonwillison.net/2025/Jul/11/kimi-k2/
Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
31–40 of 194 posts
Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#32Earlier quoted context omitted.
It does not have to be VRAM, it could be system RAM, or weights streamed from SSD storage. Reportedly, the latter method achieves around 1 token per second on computers with 64 GB of system RAM. R1 (and K2) is MoE, whereas Llama 3 is a dense model family. MoE actually makes these models practical to run on cheaper hardware. DeepSeek R1 is more comfortable for me than Llama 3 70B for exactly that reason - if it spills…
The amount of people who will be using it at 1 token/sec because there's no better option, and have 64 GB of RAM, is vanishingly small. IMHO it sets the local LLM community back when we lean on extreme quantization & streaming weights from disk to say something is possible*, because when people try it out, it turns out it's an awful experience. * the implication being, anything is possible in that scenario
Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#33Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#34Is there any way that I could do so?
Open Router? Or does kimi have their own website? Just curious to really try it out!
Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#35> 1T total / 32B active MoE model Is this the largest open-weight model?
At 1T MoE on 15.5T tokens, K2 is one of the largest open source models to date. But BAAI's TeleFM is 1T dense on 15.7T tokens: https://huggingface.co/CofeAI/Tele-FLM-1T
You can always check here: https://lifearchitect.ai/models-table/
Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#36I really really want to try this model for free since I just don't have a gpu. Is there any way that I could do so? Open Router? Or does kimi have their own website? Just curious to really try it out!
Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#37Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#38Earlier quoted context omitted.
It does not have to be VRAM, it could be system RAM, or weights streamed from SSD storage. Reportedly, the latter method achieves around 1 token per second on computers with 64 GB of system RAM. R1 (and K2) is MoE, whereas Llama 3 is a dense model family. MoE actually makes these models practical to run on cheaper hardware. DeepSeek R1 is more comfortable for me than Llama 3 70B for exactly that reason - if it spills…
The amount of people who will be using it at 1 token/sec because there's no better option, and have 64 GB of RAM, is vanishingly small. IMHO it sets the local LLM community back when we lean on extreme quantization & streaming weights from disk to say something is possible*, because when people try it out, it turns out it's an awful experience. * the implication being, anything is possible in that scenario
I will also point out that having three API-based providers deploying an impractically-large open-weights model beats the pants of having just one. Back in the day, this was called second-sourcing IIRC. With proprietary models, you're at the mercy of one corporation and their Kafkaesque ToS enforcement.
Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#39Pelican on a bicycle result: https://simonwillison.net/2025/Jul/11/kimi-k2/
Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#40Pelican on a bicycle result: https://simonwillison.net/2025/Jul/11/kimi-k2/