Kimi K3: Open Frontier Intelligence
81–90 of 1001 posts
Re: Kimi K3: Open Frontier Intelligence
#82More details: - https://platform.kimi.ai/docs/guide/kimi-k3-quickstart - https://platform.kimi.ai/docs/pricing/chat-k3 1M context, pricing is $3/$15 for 1M tokens (cache $0.3), which is extremely high for a Chinese open-weight model, but if it's truly competitive with most of the current frontier and is only behind Fable/Sol, the pricing is justified. This is 1:1 pricing of Anthropic's Sonnet series (except Sonnet 5…
Re: Kimi K3: Open Frontier Intelligence
#83Half kidding feature request for HN: Mark all AI related posts so I can filter them out, when I need a pause.
Click the link to view conversation with Kimi AI Assistant https://www.kimi.com/share/19f6b96d-fdd2-8589-8000-0000daada...
Re: Kimi K3: Open Frontier Intelligence
#84Re: Kimi K3: Open Frontier Intelligence
#85Combine with the price it will surely more costly than gpt 5.6.
Re: Kimi K3: Open Frontier Intelligence
#86This puts them on the top of the largest open models list:
Kimi K3 2.8T
DeepSeek-V4-Pro 1.6T (49B active)
Kimi K2.6 ~1T (32B active)
GLM-5.2 754B (40B active)
DeepSeek-V3.2 685B
Mistral Large 3 675B
That's one mighty large model! Moonshot is going to need the USD 500 million reportedly raised earlier this year to run this model.Re: Kimi K3: Open Frontier Intelligence
#87No blog post? Benchmarks?
Re: Kimi K3: Open Frontier Intelligence
#88Amazing to see an open source model already nearing the benchmarks of Fable and GPT 5.6 Sol! Also very cool to see LatentMoE being picked up by more models ( https://arxiv.org/abs/2601.18089 )
Re: Kimi K3: Open Frontier Intelligence
#89Say what you want about these Chinese models but they sure create competition and urgency in the space.
Re: Kimi K3: Open Frontier Intelligence
#90> We also further increased the sparsity of the Mixture of Experts (MoE): with the Stable LatentMoE framework, the model efficiently activates 16 out of 896 experts. Together with improvements in training methodology and data recipes, these structural advances give K3 roughly 2.5x the overall scaling efficiency of K2, converting compute into capability more effectively. Assuming experts are uniformly distributed (I’m…
2.5x the scaling efficiency, so 4 times the price? What is happening here? Did the subsidies dry up with the discrepancy between chinese and US models?
Kind of like scaling your personal automobile to the weight of a semi, the semi is still going to be far more efficient in moving cargo, not that the semi will cost the same to operate as the original car.