Live data from Hacker News

Kimi K2.7-Code: open-source coding model with better token efficiency

huggingface.co

201–210 of 254 posts

Re: Kimi K2.7-Code: open-source coding model with better token efficiency

#201
post #178

Output tokens are almost 5x more expensive than mimov2.5 pro/dsv4pro. I’m curious to see if Kimik2.7 is that much better. Feels like kimi are positioning themselves as the premium open source models

I find that I don't use a ton of output tokens. I'm usually around 95% cached input, 4% input, and 1% output. For me, the big thing with MiMo-V2.5-Pro and DeepSeek V4-Pro is that cached inputs are practically free. Kimi K2.7 Code is 53x more expensive for cached inputs which is 95% of my costs. If I use 95M cached input tokens, 4M input tokens, and 1M output tokens, that'd be: $18 for cached input on Kimi K2.7 Code v…

95/4/1 holds here too

Re: Kimi K2.7-Code: open-source coding model with better token efficiency

#202
post #10

I would really love to know if anyone has any experience with something like opencode + Kimi K2.6/2.7 now compared to Claude Code. What is better, what is worse, what is the cost comparison. I am currently paying $100 for the 5x Max plan, but Fable is running through the usage limits quite drastically and I cannot really say it's night and day compared to Opus. Also, I use this mostly for my side projects, so the $10…

The best is GLM (though it's not as cheap as DeepSeek or Kimi) and use it with Claude Code.

Re: Kimi K2.7-Code: open-source coding model with better token efficiency

#203
post #94

Earlier quoted context omitted.

I've kind of given up on the routers for "free" inference, as you would expect, they tend to give you sub-par thinking because they are obviously trying to conserve as much inference as possible. I've had some success turning my macbook M1 pro into a heating pad with Qwen 3.6 35B A3B MTP. Trying to use Gemini models "locally" resulted in a similar "short shrift" of effort resulting in mistakes and lots of turns. The…

> I've kind of given up on the routers for "free" inference, as you would expect, they tend to give you sub-par thinking because they are obviously trying to conserve as much inference as possible. Xiaomi MiMo ($6/mo: https://platform.xiaomimimo.com/token-plan ) & Alibaba Qwen ($50/mo: https://www.alibabacloud.com/en/campaign/ai-scene-coding ) have generous limits on fixed subscriptions.

So does Opencode Go ($10/mo: https://opencode.ai/go) for DeepSeek v4 Flash and MiMo 2.5.

Re: Kimi K2.7-Code: open-source coding model with better token efficiency

#204
post #2

I was wondering how does Anthropic and likes keep competitive when Opus is ($5 / $25) 5x times more expensive compared to Kimi K2.6 ($0.7 / $3.4) or other Chinese models, while being only marginally better. My theory is that US enterprise just can't send data to Chinese and that's understandable, but is that "the moat"?

API token price is one thing, but subscriptions on Claude are a good value. Weirdly everyone says that Claude subscriptions are subsidized because of the API price, even though (1) no one actually knows Claude's cost of inference, and (2) Chinese providers are also able to provide cheap inference, so why do they think Claude can't? I also wonder if Enterprises have deals for other API pricing that is not posted publi…

> no one actually knows Claude's cost of inference

There were some rumors stating that their margin is around 70%. So they could go much cheaper probably, talking inference only. The other thing is R&D cost...

Re: Kimi K2.7-Code: open-source coding model with better token efficiency

#205
post #37

Earlier quoted context omitted.

I think most people who've tried them both would tell you Anthropic's models are more than marginally better than Kimi. Kimi and the other open source models may score well on SWE-bench or whatever but the gap is noticeable IMHO once you actually try to use them.

It depends on what your task is and how precise your prompts are. Planning with fable or 4.8 and laying out the plan in step by step process and coding with mimo v2.5 pro or dsv4pro or qwen 3.7 max and doing a final review with 5.5 has worked really well for me for infra stuff.

Coding with sufficiently precise plan takes almost all real work from the implementator, doesn't it? So it's not a fair comparison...

Re: Kimi K2.7-Code: open-source coding model with better token efficiency

#206
post #51
post #39

Earlier quoted context omitted.

I do have this experience. I've used Claude Code (with Opus mostly), and then switched to opencode (mostly with Kimi 2.6) for my personal projects; it's based on a couple months of use. Claude Code is better. But Opencode + kimi 2.6 is workable, which is big. For bare code writing, if you know what exactly you want, most popular models are fine (deepseek, kimi, etc), it feels more or less the same as anthropic models…

>At the same time, Opus seems to understand my intent way better than e.g. deepseek. I need to be much more precise with my prompts when using deepseek - it often goes in a wrong direction if I'm lazy. This results in a workflow which feels quite a lot different from Claude Code. how much of that is Opus injecting prior conversations from memory?

I'm using Claude code + (a patched) litellm proxy + openrouter + Qwen 3.7 max/kimi k2.6/deepseek v4 pro. The only feature that doesn't work is webfetch and web search, which I've replaced with the ddg MCP. Memory, caching, and everything else works fine.

Qwen comes close to opus for planning but fable is clearly superior. Kimi and deepseek are pretty much indistinguishable from opus for coding if opus writes the plan.

I'm now testing out fable for research and planning and deepseek v4 flash for coding. I'm guessing results will be pretty similar to opus + deepseek v4 pro and costs should be lower overall.

Re: Kimi K2.7-Code: open-source coding model with better token efficiency

#207
post #10

I would really love to know if anyone has any experience with something like opencode + Kimi K2.6/2.7 now compared to Claude Code. What is better, what is worse, what is the cost comparison. I am currently paying $100 for the 5x Max plan, but Fable is running through the usage limits quite drastically and I cannot really say it's night and day compared to Opus. Also, I use this mostly for my side projects, so the $10…

I'm using Claude code + (a patched) litellm proxy + openrouter + Qwen 3.7 max/kimi k2.6/deepseek v4 pro. The only feature that doesn't work is webfetch and web search, which I've replaced with the ddg MCP and a web fetch/search pre hook to redirect the agent. Memory, caching, and everything else works fine.

Qwen comes close to opus for planning but fable is clearly superior. Results for kimi and deepseek are pretty much indistinguishable from opus for coding if opus writes the plan. The biggest difference is output cadence. Kimi for example thinks for a long time then quickly outputs a lot of text.

I'm now testing out fable for research and planning and deepseek v4 flash for coding. I'm guessing results will be pretty similar to opus + deepseek v4 pro and costs should be lower overall.

Re: Kimi K2.7-Code: open-source coding model with better token efficiency

#208

Earlier quoted context omitted.

> Trump cannot instruct Google to put picture of dogs on their homepage. I'm sorry, but that was a horrible example. Corporations have no obligation to donate money to the ballroom yet Google has donated millions.

>Corporations have no obligation to donate money to the ballroom yet Google has donated millions. Imagine living in a country where they have the obligation.

Which country have that? Pretty sure ballrooms for their supreme leader is an American thing.

Re: Kimi K2.7-Code: open-source coding model with better token efficiency

#209
post #38
post #11

I think there is some threshold after which "best" model doesn't matter, we are not that far from it. Fable now is really good, in a year or so, if Kimi catches up, even if Fable6 is much better, I think I will use kimi at 1/10th of the price. I said that about opus 4.5 at the time, thinking "this is so good, in 6-12 months the Chinese models will be as good and cheap, I will use them", but I was wrong.. I pay premiu…

Depending on who you are and how you use these models, we're already at this point

Exactly, for long running vibe coded stuff that I don't care about quality getting big and smart model is the only option. But for high quality changes where I need to have control and understand everything, where I do everything in small chunks - I can use basic model like Sonnet.

Re: Kimi K2.7-Code: open-source coding model with better token efficiency

#210

Earlier quoted context omitted.

I really hope we stop using the term "Chinese models". It has this air of Negative connotation. It's the equivalent of calling cars Japanese, which people used to do but now is almost entirely meaningless. You just call them Toyota, Honda, Lexus etc.

[flagged]

Should we call this the DARPA Internet then?
Post reply on HN