Live data from Hacker News

Kimi K2.5 Technical Report [pdf]

github.com

91–100 of 146 posts

Re: Kimi K2.5 Technical Report [pdf]

#91

Seems that K2.5 has lost a lot of the personality from K2 unfortunately, talks in more ChatGPT/Gemini/C-3PO style now. It's not explictly bad, I'm sure most people won't care but it was something that made it unique so it's a shame to see it go. examples to illustrate https://www.kimi.com/share/19c115d6-6402-87d5-8000-000062fec... (K2.5) https://www.kimi.com/share/19c11615-8a92-89cb-8000-000063ee6... (K2)

It's hard to judge from this particular question, but the K2.5 output looks at least marginally better AIUI, the only real problem with it is the snarky initial "That's very interesting" quip. Even then a British user would probably be fine with it.

Re: Kimi K2.5 Technical Report [pdf]

#92
Kimi K2T was good. This model is outstanding, based on the time I've had to test it (basically since it came out). It's so good at following my instructions, staying on task, and not getting context poisoned. I don't use Claude or GPT, so I can't say how good it is compared to them, but it's definitely head and shoulders above the open weight competitors

Re: Kimi K2.5 Technical Report [pdf]

#93
post #2

I've been using this model (as a coding agent) for the past few days, and it's the first time I've felt that an open source model really competes with the big labs. So far it's been able to handle most things I've thrown at it. I'm almost hesitant to say that this is as good as Opus.

Can you share how you're running it?

Been using K2.5 Thinking via Nano-GPT subscription and `nanocode run` and it's working quite nicely. No issues with Tool Calling so far.

Re: Kimi K2.5 Technical Report [pdf]

#94

Earlier quoted context omitted.

High end consumer SSDs can do closer to 15 GB/s, though only with PCI-e gen 5. On a motherboard with two m.2 slots that's potentially around 30GB/s from disk. Edit: How fast everything is depends on how much data needs to get loaded from disk which is not always everything on MoE models.

Would RAID zero help here?

Yes, RAID 0 or 1 could both work in this case to combine the disks. You would want to check the bus topology for the specific motherboard to make sure the slots aren't on the other side of a hub or something like that.

Re: Kimi K2.5 Technical Report [pdf]

#96

I wonder how K2.5 + OpenCode compares to Opus with CC. If it is close I would let go of my subscription, as probably a lot of people.

I've been drafting plans/specs in parallel with Opus and Kimi. Then asking them to review the others plan.

I still find Opus is "sharper" technically, tackles problems more completely & gets the nuance.

But man Kimi k2.5 can write. Even if I don't have a big problem description, just a bunch of specs, Kimi is there, writing good intro material, having good text that more than elaborates, that actually explains. Opus, GLM-4.7 have both complemented Kimi on it's writing.

Still mainly using my z.ai glm-4.7 subscription for the work, so I don't know how capable it really is. But I do tend to go for some Opus in sticky spots, and especially given the 9x price difference, I should try some Kimi. I wish I was set up for better parallel evaluation; feels like such a pain to get started.

Re: Kimi K2.5 Technical Report [pdf]

#97

How do people evaluate creative writing and emotional intelligence in LLMs? Most benchmarks seem to focus on reasoning or correctness, which feels orthogonal. I’ve been playing with Kimmy K 2.5 and it feels much stronger on voice and emotional grounding, but I don’t know how to measure that beyond human judgment.

I am trying! https://mafia-arena.com

I just don't have enough funding to do a ton of tests

Re: Kimi K2.5 Technical Report [pdf]

#98
post #46
post #9

I really like the agent swarm thing, is it possible to use that functionality with OpenCode or is that a Kimi CLI specific thing? Does the agent need to be aware of the capability?

Has anyone tried it and decided it's worth the cost; I've heard it's even more profligate with tokens?

Yes. https://x.com/swyx/status/2016381014483075561?s=20 it's not crazy, they cap it to 3 credits, and also YSK agent swarm is a closed source product

Would i use it a gain compared to Deep Research products elsewhere? Maybe, probably not but only bc it's hard to switch apps

Re: Kimi K2.5 Technical Report [pdf]

#99

Earlier quoted context omitted.

Well to be the devil's advocate: One is a household name that holds most of the world's silicon wafers for ransom, and the other sounds like a crypto scam. Also estimating valuation of Chinese companies is sort of nonsense when they're all effectively state owned.

There isn't a single % that is state owned in Moonshot AI. And don't start me with the "yeah but if the PRC" because it's gross when US can de facto ban and impose conditions even on European companies, let alone the control it has on US ones.

Funny because that's how us Americans feel about your European cookie banner litter and unilateral demands on privacy
Post reply on HN