Live data from Hacker News

Kimi K3: Open Frontier Intelligence

kimi.com

71–80 of 1001 posts

Re: Kimi K3: Open Frontier Intelligence

#71
post #37

> In our evaluations, Kimi K3 delivers frontier-level performance. Among the models tested, its overall intelligence ranks second only to Claude Fable 5 and GPT-5.6 Sol. For the complete benchmark results, see our tech blog. The full model weights of Kimi K3 will be released in the coming days. More details on the architecture, training, and evaluation will be published together with the Kimi K3 technical report. > K…

Where are you seeing this write up?

I copied that from https://platform.kimi.ai/docs/guide/kimi-k3-quickstart but it seems they updated the page to remove the benchmark score now.

Re: Kimi K3: Open Frontier Intelligence

#73
post #65
post #50

Earlier quoted context omitted.

> its overall intelligence ranks second only to Claude Fable 5 and GPT-5.6 Sol Pretty sure ranking “second” to two others means ranking third.

Yeah, bad wording it seems. Though a charitable interpretation is that Fable 5 and GPT 5.6 Sol are joint 1st place in the measurement.

If there are two folks standing at gold, nobody gets the silver medal.

Re: Kimi K3: Open Frontier Intelligence

#74
post #24

> We also further increased the sparsity of the Mixture of Experts (MoE): with the Stable LatentMoE framework, the model efficiently activates 16 out of 896 experts. Together with improvements in training methodology and data recipes, these structural advances give K3 roughly 2.5x the overall scaling efficiency of K2, converting compute into capability more effectively. Assuming experts are uniformly distributed (I’m…

No, you can't divide the entire size by the expert count. A lot of weights are constant for all tokens, so total active count is ((2800-(shared)/896)*16 + (shared))

Re: Kimi K3: Open Frontier Intelligence

#75
post #37

> In our evaluations, Kimi K3 delivers frontier-level performance. Among the models tested, its overall intelligence ranks second only to Claude Fable 5 and GPT-5.6 Sol. For the complete benchmark results, see our tech blog. The full model weights of Kimi K3 will be released in the coming days. More details on the architecture, training, and evaluation will be published together with the Kimi K3 technical report. > K…

> > K3 pushes the boundary of end-to-end knowledge work. On the GDPval-AA v2 leaderboard, Kimi K3 scores 1687. The benchmark evaluates AI models on real-world tasks across 44 occupations and 9 major industries; Kimi K3 ranks behind only Claude Fable 5 Max and GPT-5.6 Sol Max, and ahead of Claude Opus 4.8 Max at 1600.

This is the same benchmark where Sonnet 5 outperforms Opus 4.8 max.

Like all model releases, the benchmarks aren't going to tell the whole story. All of the open weight models come with amazing benchmark results now. It's hard to believe anything other than that the benchmarks are leaking into (or intentionally included) into training data.

Re: Kimi K3: Open Frontier Intelligence

#76
post #2

More details: - https://platform.kimi.ai/docs/guide/kimi-k3-quickstart - https://platform.kimi.ai/docs/pricing/chat-k3 1M context, pricing is $3/$15 for 1M tokens (cache $0.3), which is extremely high for a Chinese open-weight model, but if it's truly competitive with most of the current frontier and is only behind Fable/Sol, the pricing is justified. This is 1:1 pricing of Anthropic's Sonnet series (except Sonnet 5…

Are thinking models only the reasonable tradeoff vs using much larger non thinking ones because the cost of output tokens is below that of input tokens?

Re: Kimi K3: Open Frontier Intelligence

#77

Excited for the deepseek release this week (or at least they announced they'd release this week). Hopefully they also push even closer to SOTA.

That is exciting! I don't understand how DeepSeek can be so cheap with their cache pricing - ~0.003 usd / 1Mtok. 100x less than Kimi K3, or similar numbers against pretty much any other decently sized model to my knowledge. I've been using it whenever possible as even longer agent sessions cost few cents.

If you read DeepSeek's papers, you'll find a litany of architectural features that allow for a greatly reduced cache hit price by shrinking the size of the KV-cache.

Re: Kimi K3: Open Frontier Intelligence

#78
post #35
post #2

More details: - https://platform.kimi.ai/docs/guide/kimi-k3-quickstart - https://platform.kimi.ai/docs/pricing/chat-k3 1M context, pricing is $3/$15 for 1M tokens (cache $0.3), which is extremely high for a Chinese open-weight model, but if it's truly competitive with most of the current frontier and is only behind Fable/Sol, the pricing is justified. This is 1:1 pricing of Anthropic's Sonnet series (except Sonnet 5…

[flagged]

The thing is - as a European, I can choose between plague and cholera.

One has mostly been reliable, stayed peaceful towards us and is primarily concerned with their internal matters and the countries right next to it. They have long-term strategy and understanding of win-win situations.

The other one keeps threatening to invade/steal Greenland. Keeps waging an economic war against the entire bloc. Positions their propagandists right in our middle and does the best to influence our elections. Exports fascism and finances antidemocratic forces. Supports the genocide in that certain country. And still have their soldiers in our country, against the wishes of a majority of the population. Oh and they don't honor any treaties if they feel like it.

Easy choice.

Does that make china an angel? Hell no, they are still committed to enslaving the Uyghur people, keep threatening neighbors and are mostly han supremacists. Human rights are seen as merely a suggestion by them.

But at the time being, one is clearly more reliable than the other. Long-term, I'd like to avoid both the US and China.

Re: Kimi K3: Open Frontier Intelligence

#79
post #35

Earlier quoted context omitted.

[flagged]

Right at this moment, there are more people in the world on the side of China than on the side of the USA. Which can translate into raw market numbers at some point. So these comments are kinda moot.

[flagged]
Post reply on HN