Live data from Hacker News

Kimi K3 is not cheap

alexinch.com

21–30 of 31 posts

Re: Kimi K3 is not cheap

#22
post #20
post #9

This article seems premature to post. Right now, the price is arbitrarily set by a single provider. Why wouldn't Moonshot collect extra revenue during this exclusivity period when they knew there would be hype? The model weights are supposed to release tomorrow. Over the next several weeks, I would expect competition among open weight providers to drive down the cost, as I've seen happen with other open weight model…

That's a very fair critique. I don't mean to imply that Kimi is not at all cheaper than U.S frontier models. I more wrote this because I believe - since Chinese LLMs entered the public consciousness via DeepSeek R1, which was genuinely ~20x cheaper than o1 - there's a bit of a halo effect around Chinese models which causes people to overestimate the scale of the discount. And relative to that price anchor, Kimi is le…

[deleted]

Re: Kimi K3 is not cheap

#23

Earlier quoted context omitted.

I mean at this point their very existence depends on it so I’m not sure if I’d be surprised

Not to mention - If you switch the view to "coding tasks" on this website: Kimi K3: $3.18 per task GLM 5.2: $6.51 per task GPT 5.6 Sol: $7.02 per task Opus 5: 8.23 per task Fable: 11.70 per task So it's pretty dang cheap lol. Nobody is using frontier inference for "office tasks".

Agreed, Kimi is cheaper for coding - I say that explicitly in the post too. However I'd have to disagree with you on the "office task" front.

General office work is one of the big frontiers the labs are pushing on, and it's part of how they're justifying the value proposition to enterprise customers. It's also accounts for a big portion of the spend on RL; tasks/environments designed to train agents to navigate Slack or Salesforce. If you're Anthropic pitching Claude to a bank (taking an example I'm familiar with), coding probably accounts for ~20% tops of the workforce, and it doesn't drive direct revenues. The 'agentic coding bump', but for all your analysts, traders, and wealth managers, would be a much more attractive prospect.

I don't disagree that coding is the most successful use case so far (and probably more relevant to a HN audience). But I think the future of the labs is also contingent on them making progress on more general white collar work. I suspect that's why the Opus 5 release blog lists 3 coding benchmarks (FrontierBench, DeepSWE and FrontierCode) to 3 or 4 more general ones applicable to office work - depending on how you slice it (GDPVal, AutomationBench, Legal Agent Benchmark, BrowseComp).

Re: Kimi K3 is not cheap

#24
post #20
post #9

This article seems premature to post. Right now, the price is arbitrarily set by a single provider. Why wouldn't Moonshot collect extra revenue during this exclusivity period when they knew there would be hype? The model weights are supposed to release tomorrow. Over the next several weeks, I would expect competition among open weight providers to drive down the cost, as I've seen happen with other open weight model…

That's a very fair critique. I don't mean to imply that Kimi is not at all cheaper than U.S frontier models. I more wrote this because I believe - since Chinese LLMs entered the public consciousness via DeepSeek R1, which was genuinely ~20x cheaper than o1 - there's a bit of a halo effect around Chinese models which causes people to overestimate the scale of the discount. And relative to that price anchor, Kimi is le…

But if you compare to Anthropic's models? The cost difference is huge. Anthropic is clearly concerned that people are realizing they are expensive, since the Opus 5 blog post dedicated a lot of time to talking about how cheap the model was compared to the competition... but this doesn't hold water when I haven't seen any independent benchmarks claiming Opus 5 is cheaper than GPT-5.6-Sol, even if it is supposedly closer.

GPT-5.6-Sol is pretty competitively priced, but not all American frontier models are, and even 10% to 30% is still significant for any commodity that's as fungible as frontier models often are.

> as you say that could go down to 20-30% cheaper

I never said anything about 20% to 30%. We don't know how much it actually costs to host this model yet, and that will determine the final price. It could be just a little less, or it could be a lot less.

> once you account for quantisation

There will be no need to account for quantization. Kimi models have been 4-bit only since at least K2.5. They don't release or serve models in higher precision than that. This isn't one of those situations where LLM inference providers are debating between serving 16-bit, 8-bit, or 4-bit, and I have never seen a publicly hosted, paid model that was hosted in less than 4-bit, even if hobbyists will use sub-4-bit quantizations sometimes locally.

Re: Kimi K3 is not cheap

#25
post #23

Earlier quoted context omitted.

Not to mention - If you switch the view to "coding tasks" on this website: Kimi K3: $3.18 per task GLM 5.2: $6.51 per task GPT 5.6 Sol: $7.02 per task Opus 5: 8.23 per task Fable: 11.70 per task So it's pretty dang cheap lol. Nobody is using frontier inference for "office tasks".

Agreed, Kimi is cheaper for coding - I say that explicitly in the post too. However I'd have to disagree with you on the "office task" front. General office work is one of the big frontiers the labs are pushing on, and it's part of how they're justifying the value proposition to enterprise customers. It's also accounts for a big portion of the spend on RL; tasks/environments designed to train agents to navigate Slack…

Okay, I think that's fair, but I'm not convinced there's anybody actually doing large amounts of compute on office tasks? Do you know anybody? Can you point to anybody publicly documenting this? Can you can you point to any specific workflows where fable is being used in lieu of more basic models?

Even if you provide exceptional answers for all of these I still think it is disingenuous at best to ignore coding tasks in writing this. I have to assume coding is 90% of the use cases for the frontier.

You've made a case that the labs need these customers. You haven't made a case that the labs have these customers.

Re: Kimi K3 is not cheap

#26
post #20
post #9

This article seems premature to post. Right now, the price is arbitrarily set by a single provider. Why wouldn't Moonshot collect extra revenue during this exclusivity period when they knew there would be hype? The model weights are supposed to release tomorrow. Over the next several weeks, I would expect competition among open weight providers to drive down the cost, as I've seen happen with other open weight model…

That's a very fair critique. I don't mean to imply that Kimi is not at all cheaper than U.S frontier models. I more wrote this because I believe - since Chinese LLMs entered the public consciousness via DeepSeek R1, which was genuinely ~20x cheaper than o1 - there's a bit of a halo effect around Chinese models which causes people to overestimate the scale of the discount. And relative to that price anchor, Kimi is le…

Sir, it's 3X cheaper on coding. 3X cheaper in any industry is earth shattering. 10% is significant. 3X is really big.

Re: Kimi K3 is not cheap

#27
post #23

Earlier quoted context omitted.

Agreed, Kimi is cheaper for coding - I say that explicitly in the post too. However I'd have to disagree with you on the "office task" front. General office work is one of the big frontiers the labs are pushing on, and it's part of how they're justifying the value proposition to enterprise customers. It's also accounts for a big portion of the spend on RL; tasks/environments designed to train agents to navigate Slack…

Okay, I think that's fair, but I'm not convinced there's anybody actually doing large amounts of compute on office tasks? Do you know anybody? Can you point to anybody publicly documenting this? Can you can you point to any specific workflows where fable is being used in lieu of more basic models? Even if you provide exceptional answers for all of these I still think it is disingenuous at best to ignore coding tasks…

No I think you're right that the amount of compute spent on office work is lower than coding - although I don't have any sense for the right share. The best source I could find was an OpenAI report [1] which mentions that ~64% of enterprise token generation is via Codex, which I would expect to skew entirely towards coding. But it's hard to say how the remainder is split, what proportion is 'frontier', or whether it's representative for Anthropic.

On your questions - I've spoken to a number of execs and seniors behind closed doors but nothing public I can point to. Anecdotally, I've spoken to senior leaders at banks spending billions of tokens on one-off tasks like prepping execs for earnings calls or piloting end-to-end agent workflows for specific use cases (but mostly piecemeal/one-off).

On Fable, financial analysts I know are using it to produce research docs, models and decks - I hear that it's a big improvement for these tasks. This lot have been blindsided by the spend growth [2], the same as for coders in enterprise (e.g. Uber blowing annual budget in 4 months [3]), so I do think there's appetite and budget for a capable, cheaper open model - but, due to the price, Kimi does not obviously fill that role the way it might for coding. That said, I still largely agree with you on share - where coding has seen a broad deployment across software development, most of the office work stuff is still fairly piecemeal and certainly lower compute-spend.

I think it's fair to say I could've focussed on coding more rather than taking AA's benchmark distribution as representative - perhaps a more balanced title would be "Kimi K3 is not cheap across the board"? I guess there's also some ambiguity about what 'cheap' means - as I said elsewhere in this thread, I think when some people talk about the price of Chinese models, they imagine Deepseek competing with o1 for 1/20th of the price. Even though it is better priced for coding, Kimi isn't Deepseek-level cheap.

I do, however, think you could debate whether coding will remain at >50% total token usage going forwards - big enterprises are hunting for ways to get value out of LLMs, and the labs are investing a correspondingly large amount in generating demonstrations and RL environments to get the models up to par. At the end of the day, programmers make up ~5% of all white collar work. Of course, it's also possible that Chinese labs will shift focus to white collar applications now they've demonstrated a lead on coding cost efficiency, so, I mean who knows - it'll be interesting to get some detail when Anthropic IPOs.

Sorry for the long reply! Appreciate it's quite meandering...

[1] https://cdn.openai.com/pdf/5d1e1489-21c0-43e4-9d42-f87efdbf0...

[2] https://www.reuters.com/business/finance/australias-cba-flag...

[3] https://fortune.com/2026/05/26/uber-coo-ai-spending-tokens-c...

Re: Kimi K3 is not cheap

#28
post #20

Earlier quoted context omitted.

That's a very fair critique. I don't mean to imply that Kimi is not at all cheaper than U.S frontier models. I more wrote this because I believe - since Chinese LLMs entered the public consciousness via DeepSeek R1, which was genuinely ~20x cheaper than o1 - there's a bit of a halo effect around Chinese models which causes people to overestimate the scale of the discount. And relative to that price anchor, Kimi is le…

Sir, it's 3X cheaper on coding. 3X cheaper in any industry is earth shattering. 10% is significant. 3X is really big.

It is cheaper than GPT-5.6 and Anthropic's models. But then on coding agents specifically (if we take artificial analysis, at least - I'm sure there's a better meta-review you could do) Grok-4.5 scores better and is 20% cheaper. Of course there are other reasons you might prefer Kimi to Grok. But still, if coding is the main concern, is Kimi opening up somewhere new on the cost-capability Pareto frontier? Not necessarily.

Re: Kimi K3 is not cheap

#29

It feels like this "Kimi is a token hog" meme is 100% astroturfed by Anthropic. It's cheap. Believe your own eyes.

I mean at this point their very existence depends on it so I’m not sure if I’d be surprised

Well, does it? I will probably switch to Kimi K3. But Anthropic does have the smartest model. They just can't charge $20 per month for it anymore and reject people from around the world from using Mythos. Drop the price by at least 2 times down to $10, better down to $5 and allow me to use Mythos and I'll switch back to Anthropic. The $2 price difference is worth it: I get better smarter models, after all. But for $3 vs $20 and on top of that they don't even let me use Fable(paying additional money only) and Mythos(can't even use it at all) it's a no brainer. Anthropic just need to get real and drop the prices, it runs on the same hardware so how bad can it be for them if people can host Kimi K3 and let people use it for $3? Haven't tried Kimi K3 as I have a sponsored Claude subscription.

Re: Kimi K3 is not cheap

#30

Earlier quoted context omitted.

I mean at this point their very existence depends on it so I’m not sure if I’d be surprised

Well, does it? I will probably switch to Kimi K3. But Anthropic does have the smartest model. They just can't charge $20 per month for it anymore and reject people from around the world from using Mythos. Drop the price by at least 2 times down to $10, better down to $5 and allow me to use Mythos and I'll switch back to Anthropic. The $2 price difference is worth it: I get better smarter models, after all. But for $3…

You are completely off base with your assumption about what it costs to serve models. Nobody is gonna serve you K3 as cheaply as Anthropic is serving you Opus on any of their plans. Nobody has that kind of money to burn. Kimi does not even sell you a plan right now.
Post reply on HN