Seems that K2.5 has lost a lot of the personality from K2 unfortunately, talks in more ChatGPT/Gemini/C-3PO style now. It's not explictly bad, I'm sure most people won't care but it was something that made it unique so it's a shame to see it go. examples to illustrate https://www.kimi.com/share/19c115d6-6402-87d5-8000-000062fec... (K2.5) https://www.kimi.com/share/19c11615-8a92-89cb-8000-000063ee6... (K2)
Kimi K2.5 Technical Report [pdf]
91–100 of 146 posts
Re: Kimi K2.5 Technical Report [pdf]
#92Re: Kimi K2.5 Technical Report [pdf]
#93I've been using this model (as a coding agent) for the past few days, and it's the first time I've felt that an open source model really competes with the big labs. So far it's been able to handle most things I've thrown at it. I'm almost hesitant to say that this is as good as Opus.
Can you share how you're running it?
Re: Kimi K2.5 Technical Report [pdf]
#94Earlier quoted context omitted.
High end consumer SSDs can do closer to 15 GB/s, though only with PCI-e gen 5. On a motherboard with two m.2 slots that's potentially around 30GB/s from disk. Edit: How fast everything is depends on how much data needs to get loaded from disk which is not always everything on MoE models.
Would RAID zero help here?
Re: Kimi K2.5 Technical Report [pdf]
#95It's a decent model but works best with kimi CLI, not CC or others.
Re: Kimi K2.5 Technical Report [pdf]
#96I wonder how K2.5 + OpenCode compares to Opus with CC. If it is close I would let go of my subscription, as probably a lot of people.
I still find Opus is "sharper" technically, tackles problems more completely & gets the nuance.
But man Kimi k2.5 can write. Even if I don't have a big problem description, just a bunch of specs, Kimi is there, writing good intro material, having good text that more than elaborates, that actually explains. Opus, GLM-4.7 have both complemented Kimi on it's writing.
Still mainly using my z.ai glm-4.7 subscription for the work, so I don't know how capable it really is. But I do tend to go for some Opus in sticky spots, and especially given the 9x price difference, I should try some Kimi. I wish I was set up for better parallel evaluation; feels like such a pain to get started.
Re: Kimi K2.5 Technical Report [pdf]
#97How do people evaluate creative writing and emotional intelligence in LLMs? Most benchmarks seem to focus on reasoning or correctness, which feels orthogonal. I’ve been playing with Kimmy K 2.5 and it feels much stronger on voice and emotional grounding, but I don’t know how to measure that beyond human judgment.
I just don't have enough funding to do a ton of tests
Re: Kimi K2.5 Technical Report [pdf]
#98I really like the agent swarm thing, is it possible to use that functionality with OpenCode or is that a Kimi CLI specific thing? Does the agent need to be aware of the capability?
Has anyone tried it and decided it's worth the cost; I've heard it's even more profligate with tokens?
Would i use it a gain compared to Deep Research products elsewhere? Maybe, probably not but only bc it's hard to switch apps
Re: Kimi K2.5 Technical Report [pdf]
#99Earlier quoted context omitted.
Well to be the devil's advocate: One is a household name that holds most of the world's silicon wafers for ransom, and the other sounds like a crypto scam. Also estimating valuation of Chinese companies is sort of nonsense when they're all effectively state owned.
There isn't a single % that is state owned in Moonshot AI. And don't start me with the "yeah but if the PRC" because it's gross when US can de facto ban and impose conditions even on European companies, let alone the control it has on US ones.