Live data from Hacker News

Kimi K3: Open Frontier Intelligence

kimi.com

331–340 of 1001 posts

Re: Kimi K3: Open Frontier Intelligence

#331

Very interesting to see how Gemini 3.5 Pro stacks up against this new wave of models. Hope they have something similar to a Gemini 3.1 moment soon. Their speciality has always been math and multi modal intelligence and the new models are recently all very coding focused.

Why Gemini 3.5 Pro in particular?

The only major player left in this round if I’m not mistaken.

Re: Kimi K3: Open Frontier Intelligence

#332
post #273
post #246

Earlier quoted context omitted.

They will release the full weights by 7/27 along with support in vLLM. Source: their release blog on WeChat. https://mp.weixin.qq.com/s/V4xhEIy8xDXSMDPrPkmUAQ

>We are currently working closely with our inference partners and open-source maintainers to align the technical details and ensure the model can be reliably deployed across the ecosystem. The full model weights will be released by July 27, 2026. Further details regarding the architecture, training, and evaluation will be released with the Kimi K3 technical report. (translated by chrome) 11 days is a long time. It do…

Actually it does for a massive model, serving it correctly is not easy.

I believe Kimi also does some sort of Q&A and eval for day 0 partners, since early on a long of inference providers just weren’t running their models properly.

Re: Kimi K3: Open Frontier Intelligence

#333
post #37

> In our evaluations, Kimi K3 delivers frontier-level performance. Among the models tested, its overall intelligence ranks second only to Claude Fable 5 and GPT-5.6 Sol. For the complete benchmark results, see our tech blog. The full model weights of Kimi K3 will be released in the coming days. More details on the architecture, training, and evaluation will be published together with the Kimi K3 technical report. > K…

> Maybe another DeepSeek moment right here. Surely not... What made DeepSeek disruptive was that the cost was 10X lower. In this case, the cost is about 2X lower the Sol I think? At 2X, you're pretty close to the error margins due to token efficiency etc... I'd say this is "on trend" for open models catching up to frontier labs, but its not a "change in the trend" like DeepSeek was IMO.

It was also disruptive because it was open weight, meaning anyone and their dog could theoretically compete with the frontier labs for their inference revenue.

The frontier labs need to recoup a huge amount of cash to cover their model development costs, and justify their valuations. That’s plausible when they’re only ones capable of selling inference on these models, it a lot less plausible when models themselves become cheap commodities, and you’re just competing on your ability to provide compute. Anthropic and OpenAI can’t compete with people like AWS on that front.

Re: Kimi K3: Open Frontier Intelligence

#334
post #227

Some official benchmark numbers posted in Chinese social media (I am sure they will publish an English blogpost later too): https://mp.weixin.qq.com/s/V4xhEIy8xDXSMDPrPkmUAQ Generally looks like a Sol/Fable tier model, better across the board than Opus 4.8. (Edit) English blogpost is up now: https://www.kimi.com/blog/kimi-k3

[flagged]

Re: Kimi K3: Open Frontier Intelligence

#335

Very interesting to see how Gemini 3.5 Pro stacks up against this new wave of models. Hope they have something similar to a Gemini 3.1 moment soon. Their speciality has always been math and multi modal intelligence and the new models are recently all very coding focused.

Bloomberg has an exclusive today about how internal metrics on Gemini 3.5 Pro are not good enough, thus the release is delayed. (Not posting link coz paywall)

https://www.reuters.com/business/google-gemini-launch-delaye...

Re: Kimi K3: Open Frontier Intelligence

#336
post #2

More details: - https://platform.kimi.ai/docs/guide/kimi-k3-quickstart - https://platform.kimi.ai/docs/pricing/chat-k3 1M context, pricing is $3/$15 for 1M tokens (cache $0.3), which is extremely high for a Chinese open-weight model, but if it's truly competitive with most of the current frontier and is only behind Fable/Sol, the pricing is justified. This is 1:1 pricing of Anthropic's Sonnet series (except Sonnet 5…

I've been avidly using Fable since it was re-released and while it has been excellent at building the apps I want, the reasoning has been completely opaque.

Kim, however, has exposed the whole reasoning trace, or enough of it to matter. I'd almost forgotten how nice it is to see this. I've been able to see all of the weird twist and turns it takes and it is joyful. But also, far, far more informative and means I can debug ideas far more thoroughly. Also, at a first glance it seems to have gotten quite far on a niche hobby horse of mine that no LLM has been able to crack. I'll be testing this more for sure.

Re: Kimi K3: Open Frontier Intelligence

#337

Earlier quoted context omitted.

[flagged]

Is this really true? I was led to believe my company had an enterprise zero data retention agreement with them and it’s why we didn’t get access to Fable Is there proof of what you’re saying or is it just a guess?

[deleted]

Re: Kimi K3: Open Frontier Intelligence

#338

Earlier quoted context omitted.

It's like reading Anthropic's obituary.

This is weird and reactionary. Lots of organizations are continuing to refuse to use chinese models due to security and IP concerns. Anthropic/american models aren't going anywhere anytime soon.

Cursor will rebrand it as Composer 3.0 to assuage any such concerns, as they did with the previous Kimi models.

Re: Kimi K3: Open Frontier Intelligence

#340

Another deepseek moment? it seems they have fully caught with fable tier of models, and this was a lot sooner than was expected.

Yeah, I would have expected Zhipu to ship a Fable-adjacent model by the end of the year, but the jump from Kimi 2.7 (which I think is just barely at the level where it is genuinely helpful for coding) to this is absolutely bonkers. And this is clearly not just benchmaxing; this thing actually works. If you told me I could only use this and never use Fable or Sol again, I'd shrug and not feel like I'd lost much.

Now it seems the best way to tell if a frontier model is benchmaxxed is to check if it can autonomously solve a major open mathematical problem.
Post reply on HN