Live data from Hacker News

Kimi K2.5 Technical Report [pdf]

github.com

111–120 of 146 posts

Re: Kimi K2.5 Technical Report [pdf]

#111
I tried this today. It's good - but it was significantly less focused and reliable than Opus 4.5 at implementing some mostly-fleshed-out specs I had lying around for some needed modifications to an enterprise TS node/express service. I was a bit disappointed tbh, the speed via fireworks.ai is great, they're doing great work on the hosting side. But I found the model had to double-back to fix type issues, broken tests, etc, far more than Opus 4.5 which churned through the tasks with almost zero errors. In fact, I gave the resulting code to Opus, simply said it looked "sloppy" and Opus cleaned it up very quickly.

Re: Kimi K2.5 Technical Report [pdf]

#112

Seems that K2.5 has lost a lot of the personality from K2 unfortunately, talks in more ChatGPT/Gemini/C-3PO style now. It's not explictly bad, I'm sure most people won't care but it was something that made it unique so it's a shame to see it go. examples to illustrate https://www.kimi.com/share/19c115d6-6402-87d5-8000-000062fec... (K2.5) https://www.kimi.com/share/19c11615-8a92-89cb-8000-000063ee6... (K2)

[flagged]

Disagree, i've found kimi useful in solving creative coding problems gemini, claude, chatgpt etc failed at. Or, it is far better at verifying, augmenting and adding to human reviews of resumes for positions. It catches missed detials humans and other llm's routinley miss. There is something special to K2.

Re: Kimi K2.5 Technical Report [pdf]

#114

Earlier quoted context omitted.

Well to be the devil's advocate: One is a household name that holds most of the world's silicon wafers for ransom, and the other sounds like a crypto scam. Also estimating valuation of Chinese companies is sort of nonsense when they're all effectively state owned.

There isn't a single % that is state owned in Moonshot AI. And don't start me with the "yeah but if the PRC" because it's gross when US can de facto ban and impose conditions even on European companies, let alone the control it has on US ones.

I'm not sure if that is accurate, most of the funding they've got is from Tencent and Alibaba, and we know what happened to Jack Ma the second he went against the party line. These two are defacto state owned enterprises. Moonshot is unlikely to be for sale in any meaningful way so its valuation is moot.

[0] https://en.wikipedia.org/wiki/Moonshot_AI#Funding_and_invest...

Re: Kimi K2.5 Technical Report [pdf]

#115

Earlier quoted context omitted.

Did you use Kimi Code or some other harness? I used it with OpenCode and it was bumbling around through some tasks that Claude handles with ease.

Are you on the latest version? They pushed an update yesterday that greatly improved Kimi K2.5’s performance. It’s also free for a week in OpenCode, sponsored by their inference provider

But it may be a quantized model for the free version.

Re: Kimi K2.5 Technical Report [pdf]

#116
post #18
post #14

Earlier quoted context omitted.

I'm not running it locally (it's gigantic!) I'm using the API at https://platform.moonshot.ai

Just curious - how does it compare to GLM 4.7? Ever since they gave the $28/year deal, I've been using it for personal projects and am very happy with it (via opencode). https://z.ai/subscribe

Kimi k2.5 is a beast, speaks very human like (k2 was also good at this) and completes whatever I throw at it. However, the glm quarterly coding plan is too good of a deal. The Christmas deal ends today, so I’d still suggest to stick to it. There will always come a better model.

Re: Kimi K2.5 Technical Report [pdf]

#118

How do people evaluate creative writing and emotional intelligence in LLMs? Most benchmarks seem to focus on reasoning or correctness, which feels orthogonal. I’ve been playing with Kimmy K 2.5 and it feels much stronger on voice and emotional grounding, but I don’t know how to measure that beyond human judgment.

https://eqbench.com/index.html

Re: Kimi K2.5 Technical Report [pdf]

#120
post #14
post #5

Earlier quoted context omitted.

Out of curiosity, what kind of specs do you have (GPU / RAM)? I saw the requirements and it's a beyond my budget so I am "stuck" with smaller Qwen coders.

I'm not running it locally (it's gigantic!) I'm using the API at https://platform.moonshot.ai

What's the point of using an open source model if you're not self-hosting?
Post reply on HN