I tried Kimi on a few coding problems that Claude was spinning on. It’s good. It’s huge, way too big to be a “local” model — I think you need something like 16 H200s to run it - but it has a slightly different vibe than some of the other models. I liked it. It would definitely be useful in ensemble use cases at the very least.
Reasonable speeds are possible with 4bit quants on 2 512GB Mac Studios (MLX TB4 Ring - see https://x.com/awnihannun/status/1943723599971443134 ) or even a single socket Epyc system with >1TB of RAM (about the same real world memory throughput as the M Ultra). So $20k-ish to play with it. For real-world speeds though yeah, you'd need serious hardware. This is more of a "deploy your own stamp" model, less a "local" mod…
Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
151–160 of 194 posts
Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#152Earlier quoted context omitted.
Someone at openai did say it was too big to host at home, so you could be right. They will probably be benchmaxxing, right now, searching for a few evals they can beat.
These are all "too big to host at home". I don't think that is the issue here. https://github.com/MoonshotAI/Kimi-K2/blob/main/docs/deploy_... "The smallest deployment unit for Kimi-K2 FP8 weights with 128k seqlen on mainstream H200 or H20 platform is a cluster with 16 GPUs with either Tensor Parallel (TP) or "data parallel + expert parallel" (DP+EP)." 16 GPUs costing ~$30k each. No one is running a ~$500k server at…
Not sure if they’ll trust a Chinese model but dropping $50-100k for a quantized model that replaces, say, 10 paralegals is good enough for a law firm
Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#153Earlier quoted context omitted.
These are all "too big to host at home". I don't think that is the issue here. https://github.com/MoonshotAI/Kimi-K2/blob/main/docs/deploy_... "The smallest deployment unit for Kimi-K2 FP8 weights with 128k seqlen on mainstream H200 or H20 platform is a cluster with 16 GPUs with either Tensor Parallel (TP) or "data parallel + expert parallel" (DP+EP)." 16 GPUs costing ~$30k each. No one is running a ~$500k server at…
The real users for these open source models are businesses that want something on premises for data privacy reasons Not sure if they’ll trust a Chinese model but dropping $50-100k for a quantized model that replaces, say, 10 paralegals is good enough for a law firm
Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#154Earlier quoted context omitted.
Reasonable speeds are possible if you pay someone else to run it. Right now both NovitaAI and Parasail are running it, both available through Openrouter and both promising not to store any data. I'm sure the other big model hosters will follow if there's demand. I may not be able to reasonably run it myself, but at least I can choose who I trust to run it and can have inference pricing determined by a competitive mar…
I’m actually finding Claude 4 Sonnet’s thinking model to be too slow to meet my needs. It literally takes several minutes per query on Cursor. So running it locally is the exact opposite of what I’m looking for. Rather, I’m willing to pay more, to have it be run on a faster than normal cloud inference machine. Anthropic is already too slow. Since this model is open source, maybe someone could offer it at a “premium”…
Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#155So far, I like the answer quality and its voice (a bit less obsequious than either ChatGPT or DeepSeek, more direct), but it seems to badly mangle the format of its answers more often than I've seen with SOTA models (I'd include DeepSeek in that category, or close enough).
Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#156Earlier quoted context omitted.
This is a dumb question I know, but how expensive is model distillation? How much training hardware do you need to take something like this and create a 7B and 12B version for consumer hardware?
The process involves running the original model. You can rent these big GPUs for ~$10 per hour, so that is ~$160 per hour for as long as it takes
Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#157Technical strengths aside, I’ve been impressed with how non-robotic Kimi K2 is. Its personality is closer to Anthropic’s best: pleasant, sharp, and eloquent. A small victory over botslop prose.
Interestingly enough, EQ-Bench/Creative Writing Bench doesn't spot this despite clearly having it in their samples. This makes me trust it even less.
Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#158All the AI models are no using em-dashes. ChatGPT keeps using them even after explicitly told not to. Anybody know what’s up with these models?
I don't know, but as someone who likes using em-dashes in my writing it is disappointing that they have become a marker of LLM slop.
Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#159I tried Kimi on a few coding problems that Claude was spinning on. It’s good. It’s huge, way too big to be a “local” model — I think you need something like 16 H200s to run it - but it has a slightly different vibe than some of the other models. I liked it. It would definitely be useful in ensemble use cases at the very least.
All the models I tested, which includes Sonnet 4, DeepSeekR1, 4o and Gemini 2.5 understand this isn't your normal email and what I ask is literary/philosophical criticism, not remarks about conventions, formatting or how to convey my message in a more impactful way.
Quick fix:
- Replace half the commas with periods.
- Let one metaphor bloom and then stop.
- Read it aloud; wherever you inhale, insert a line break.
Before:
>And if the collective imagination is so seized by it, it's because
we are already—not in some other world—but in this very life, in a
state of socio-, photo-, and videosynthesis.
After:
>The collective imagination is seized.
>We are not in another world.
>We are in this life, in a state of videosynthesis.
Yes, this quote is by Baudrillard. None of the other models fixated on the fact that it’s an email (I only used the word once in the prompt). My gut feeling is that this reflects not so much a lack of intelligence as a difference in model personality. Here's what it replied when I shared Gemini's analysis. The point was to have it understand that: Gemini, second sentence: It is less a simple email than an intellectual
and personal exploration still in the making.
Point deflected as soon as understood: **Gemini is polite and pedagogical; I’m going to be blunt.**
1. **This isn’t an email, it’s a trip.**
...
2. **The content is good, the format is not.**
...
3. **Gemini is right about one thing:**
...
* cut 80% of the material,
* ask **one clear question**
* put the rest in a separate Google Doc or PDF for posterity.
4. **Brutal summary:**
Keep your draft for yourself — it’s valuable. But for the email, write a new one in **three sentences**:
* thank them,
* recall the context in one line,
* ask the question.
You’ve already got Gemini’s answer to stroke your ego — now give me the short version you’re actually going to send.
The solution to sycophancy is not disparagement (misplaced criticism). The classical true/false positive/negative dilemma is at play here. I guess the bot got caught in the crossfire of 1°) its no-bullshit attitude (it can only be an attitude) 2°) preference for delivering blunt criticism over insincere flattery 3°) being a helpful assistant. Remove point 3°), and it could have replied: "I'm not engaging in this nonsense". Preserve it and it will politely suggest that you condense your bullshit text, because shorter explanations are better than long winding rants (it's probably in the prompt).Re: Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
#160So far, I like the answer quality and its voice (a bit less obsequious than either ChatGPT or DeepSeek, more direct), but it seems to badly mangle the format of its answers more often than I've seen with SOTA models (I'd include DeepSeek in that category, or close enough).
Which host did you use? I noticed the same using parasail. Switching to novita and temp 0.4 solved it.