Live data from Hacker News

Kimi K3: Open Frontier Intelligence

kimi.com

81–90 of 1001 posts

Re: Kimi K3: Open Frontier Intelligence

#82
post #2

More details: - https://platform.kimi.ai/docs/guide/kimi-k3-quickstart - https://platform.kimi.ai/docs/pricing/chat-k3 1M context, pricing is $3/$15 for 1M tokens (cache $0.3), which is extremely high for a Chinese open-weight model, but if it's truly competitive with most of the current frontier and is only behind Fable/Sol, the pricing is justified. This is 1:1 pricing of Anthropic's Sonnet series (except Sonnet 5…

It seems the subsidized era is nearing its end and we'll see a convergence on API pricing before a pulling of subscriptions pricing.

Re: Kimi K3: Open Frontier Intelligence

#83
post #5

Half kidding feature request for HN: Mark all AI related posts so I can filter them out, when I need a pause.

I think we have a need to revise the old let me Google that for you thing

Click the link to view conversation with Kimi AI Assistant https://www.kimi.com/share/19f6b96d-fdd2-8589-8000-0000daada...

Re: Kimi K3: Open Frontier Intelligence

#85
Not worth it. I have just tried a single prompt in the web interface and it is still not finish reasoning. It thinks too much and often repeats the same stuff over and over.

Combine with the price it will surely more costly than gpt 5.6.

Re: Kimi K3: Open Frontier Intelligence

#86
> Kimi K3 is Kimi’s most capable model to date, with 2.8 trillion parameters.

This puts them on the top of the largest open models list:

  Kimi K3            2.8T
  DeepSeek-V4-Pro    1.6T (49B active)
  Kimi K2.6          ~1T (32B active)
  GLM-5.2            754B (40B active)
  DeepSeek-V3.2      685B
  Mistral Large 3    675B
That's one mighty large model! Moonshot is going to need the USD 500 million reportedly raised earlier this year to run this model.

Re: Kimi K3: Open Frontier Intelligence

#90
post #32
post #24

> We also further increased the sparsity of the Mixture of Experts (MoE): with the Stable LatentMoE framework, the model efficiently activates 16 out of 896 experts. Together with improvements in training methodology and data recipes, these structural advances give K3 roughly 2.5x the overall scaling efficiency of K2, converting compute into capability more effectively. Assuming experts are uniformly distributed (I’m…

2.5x the scaling efficiency, so 4 times the price? What is happening here? Did the subsidies dry up with the discrepancy between chinese and US models?

Scaling efficiency simply means if you took the first small model and scaled it up to the big model it would take 2.5x the resources to run. Not the that larger model is going to be any cheaper.

Kind of like scaling your personal automobile to the weight of a semi, the semi is still going to be far more efficient in moving cargo, not that the semi will cost the same to operate as the original car.

Post reply on HN