Live data from Hacker News

Kimi K3: Open Frontier Intelligence

kimi.com

231–240 of 1001 posts

Re: Kimi K3: Open Frontier Intelligence

#231
post #82

Earlier quoted context omitted.

It seems the subsidized era is nearing its end and we'll see a convergence on API pricing before a pulling of subscriptions pricing.

That’s not what this indicates. This is the biggest and most expensive to serve, and most capable open weights model yet. They’re just pricing it in line with capabilities. Kimi also offers generous subscriptions. Subs aren’t going anywhere. Think of subs like running an insurance business. There might be some users you lose money on (ones who max out their weekly quota without fail), but they’re managed such that th…

> They’re just pricing it in line with capabilities.

So... convergence?

> but they’re managed such that the average subscription turns a healthy profit.

It didn't work like that, or at least that's not how it played out. People max-out their subs all the time which is why strict and multiple limits were implemented by all providers. Also, I subscribe to z.ai and recently they dropped the quota significantly that now their sub offers less than Claude and OpenAI. It's still x5-6 what it would cost on API costs though.

> inference is just cheaper in terms of ops TCO than people assume, and API margins are very high.

API margins (at least american ones) are probably healthy. But I don't think that inference is that cheap. It would cost 300-500k to just run GLM 5.2. There are lots of other factors too: reliability (can you keep the GPUs running all time), electricity cost, sys. admin costs, location costs, etc.. I wouldn't be surprised if the API margins are quite close to operational costs.

Re: Kimi K3: Open Frontier Intelligence

#232

It does seem to have retained the K2 series's creative writing abilities, at least with the prompts I've tested so far.

Good that they are keeping it, Kimis way of speaking and conveying some sort of EQ is absolutely the best. The other models might be better at certain things, but nothing comes close to how good Kimi is at understanding language, emotions and reading the room in conversations.

I should maybe also mention that I have not used the later models like Opus or Fable, so my opinion might be a bit outdated.

When I remember that this site even showed Kimi having the highest score at one point https://eqbench.com

Re: Kimi K3: Open Frontier Intelligence

#233

This is too expensive to be a viable model. If it were $5/1m output, it might be another story. At these prices, there's no reason to use this over GPT 5.6.

[flagged]

Is this really true? I was led to believe my company had an enterprise zero data retention agreement with them and it’s why we didn’t get access to Fable

Is there proof of what you’re saying or is it just a guess?

Re: Kimi K3: Open Frontier Intelligence

#234

Earlier quoted context omitted.

Reuters has been reporting that Chinese government is undergoing similar investigation to the US; blocking the export of domestic frontier models. They boil down to "anonymous sources" but it does seem inevitable as the tech gets stronger and stronger.

It came (at least in part) from a document in May where the CCP pretty much said that they will need to review models to make sure they don't threaten national security. Which basically translates too "Don't give away tools that can be used to undermine your own goals".

So much for the speculation that China was encouraging the release of free/cheap models to mess with the US AI economy.

Re: Kimi K3: Open Frontier Intelligence

#235
post #131

Earlier quoted context omitted.

They've removed the paragraph about releasing model weights.

Does that mean this one won't be open source?

nitpicking and beating a dead horse, but it was never going to be open source, at best open weight.

Re: Kimi K3: Open Frontier Intelligence

#237
post #222

Earlier quoted context omitted.

[flagged]

In context it seems your recommendation is to instead send those data to models within Chinese nation-network space. I’m not here to defend US frontier model companies; your accusation is probably accurate. But I doubt sending data to China is an improvement.

with open weight models, you have three other options

A) use a provider that pinky-swears not to store your data. they obviously don't give a fuck about 'distillation attacks', so they have little motivation to voluntarily monitor and store your queries. reasonably high likelihood of privacy.

B) rent the hardware and run the model yourself. very high likelihood of privacy.

C) buy the hardware and run the model yourself. absolute certainty of privacy.

Re: Kimi K3: Open Frontier Intelligence

#238

I'm a bit nervous this one isn't going to be open-weights. Any mention of "open" has been struck from the literature for this model (it was present an hour ago). We don't even know active params? At this pricing, I'll be surprised if it's open.

[flagged]

Re: Kimi K3: Open Frontier Intelligence

#239

Any updated Pareto frontier graphs? https://paraplouis.github.io/llm-pareto-frontier/ is quite out of date now.

I generally rely on LMArena for this: https://arena.ai/leaderboard/code/webdev/pareto But it does take some days after model release before they collect enough data.

LMArena's "code" leaderboard is really skewed since it's a front-end JS code and design leaderboard. It generates a demo app with two models and then asks "do you prefer A or B". People can look at the code, but most of the time it's just going to be which one looks nicer.

Models that people like the design aesthetic of (Claude, GLM) tend to do better in LMArena than they do on other benchmarks. Design matters, but you look at a model like GPT-5.5 and it's behind Kimi K2.6, Sonnet 4.6, Qwen3.7 Max, and GLM-5.1 on LMArena's code leaderboard. Then you look at benchmarks like DeepSWE and GPT-5.5 blows them out of the water with only Fable and GPT-5.6 beating it.

I'm not saying that the LMArena leaderboard isn't useful, but I'm not sure how much weight I'd give it as a "code" leaderboard. I think often times it's a design comparison of simple front-end React apps rather than a coding comparison. GLM-5.2 is a very good model, but when you look at DeepSWE or Terminal-Bench v2, GPT-5.5 is well ahead.

Re: Kimi K3: Open Frontier Intelligence

#240
post #2

More details: - https://platform.kimi.ai/docs/guide/kimi-k3-quickstart - https://platform.kimi.ai/docs/pricing/chat-k3 1M context, pricing is $3/$15 for 1M tokens (cache $0.3), which is extremely high for a Chinese open-weight model, but if it's truly competitive with most of the current frontier and is only behind Fable/Sol, the pricing is justified. This is 1:1 pricing of Anthropic's Sonnet series (except Sonnet 5…

Does it have safety guardrails that constantly false positive like Claude does? The only obvious change I’ve seen since opus 4.6 came out is that it constantly flags my requests (no, I’m not doing biology research or security research, yes, it flags for both of those things).

Recently, they backported the blocks to Opus 4.8, so I’m reluctantly stuck on sonnet.

I probably could successfully apply to get special approval to use claude code unencumbered, but I don’t think it is ethical to support tooling that’s built so a central authority gets to decide what intellectual endeavors and knowledge work are permissible, and what are not.

Post reply on HN