Live data from Hacker News

Kimi K3: Open Frontier Intelligence

kimi.com

571–580 of 1001 posts

Re: Kimi K3: Open Frontier Intelligence

#571

[flagged]

There is no such thing as an unbiased opinion, opinions always contain some sort of value judgement. Besides, training data contains biases.

What is plausible is that they haven't made any attempts to explicitly steer certain opinions into a certain direction, and just let the model take over the bias of the training data, whatever that may be. Filtering in the front end is the easy, lazy way out to be legally compliant.

Re: Kimi K3: Open Frontier Intelligence

#572

My testing prompt for these models is by no means objective or repeatable (like the pelican) but it's a nice test of curiosity: > Impress me with a 1 page html file Result: https://ydaurtg3fdwhq.kimi.page/ Came out looking pretty cool! By contrast, Fable produced a moderately more interesting "live observatory" of the solar system.

Nice qualitative test!

Just like you I am super impressed by Kimi K3.

I do a qualitative benchmark series making 3D explainers and so here's Kimi K3 vs Claude Fable:

https://generative-ai.review/2026/07/kimi-k3-rush-test-vs-cl...

I've put links to the posts on GLM5.2, Opus 4.8, Chat GPT 5.5. I grab video screencaps so you can compare in detail. The full interactive Kimi output is at the bottom of the post if you want a comprehensive 3D play around

Re: Kimi K3: Open Frontier Intelligence

#573

Earlier quoted context omitted.

it also undercuts American dominance, which is something China is always happy to invest in, even if it doesn't immediately mean Chinese dominance

This. More people should read Simon Wardley. Also China's silent but powerful support of Russia and its invasion.

[deleted]

Re: Kimi K3: Open Frontier Intelligence

#574
post #373
post #111

Pelican: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... - rendered via the OpenRouter API: https://openrouter.ai/moonshotai/kimi-k3 95 input, 16,658 output = 25 cents! https://www.llm-prices.com/#it=95&ot=16658&ic=3&oc=15 (13,241 of those were reasoning tokens.) I think that's the most expensive pelican I've rendered through a Chinese model so far.

Wrote this up in a bit more detail on my blog, including some thoughts on what value the pelican benchmark can still provide here: https://simonwillison.net/2026/Jul/16/kimi-k3/

In regards to your post and the 16k reasoning output.

Try setting reasoning levels yourself manually. We see in the benchmarks that one of the graphs shows low, mid, max, so its clearly there.

I had the same issue with GLM 5.2 only offering high/max.

By playing around with openai compatible protocol, and setting the reasoning level from none, low ... high, xhigh and testing a flawed logic test.

It was easy to see that GLM had all the different reasoning levels. Low was like one line, medium did a few, high started to really expand, xhigh was a page or 2, max was MAX.

Very sure that you can force K3 into using less reasoning.

Re: Kimi K3: Open Frontier Intelligence

#576
post #111

Pelican: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... - rendered via the OpenRouter API: https://openrouter.ai/moonshotai/kimi-k3 95 input, 16,658 output = 25 cents! https://www.llm-prices.com/#it=95&ot=16658&ic=3&oc=15 (13,241 of those were reasoning tokens.) I think that's the most expensive pelican I've rendered through a Chinese model so far.

I'm still waiting for the day that one of these models interprets the request weird and outputs an SVG of a Pelican case.

https://canada.newark.com/productimages/large/en_US/4492516....

In the field I work in, if someone says "Pelican", 99.99% of the time it's going to be an equipment case. We never have reason or need to refer to the actual bird.

Re: Kimi K3: Open Frontier Intelligence

#577
post #572

My testing prompt for these models is by no means objective or repeatable (like the pelican) but it's a nice test of curiosity: > Impress me with a 1 page html file Result: https://ydaurtg3fdwhq.kimi.page/ Came out looking pretty cool! By contrast, Fable produced a moderately more interesting "live observatory" of the solar system.

Nice qualitative test! Just like you I am super impressed by Kimi K3. I do a qualitative benchmark series making 3D explainers and so here's Kimi K3 vs Claude Fable: https://generative-ai.review/2026/07/kimi-k3-rush-test-vs-cl... I've put links to the posts on GLM5.2, Opus 4.8, Chat GPT 5.5. I grab video screencaps so you can compare in detail. The full interactive Kimi output is at the bottom of the post if you want…

Very neat, this is definitely a more complete prompt than mine!

Re: Kimi K3: Open Frontier Intelligence

#578
post #377

Earlier quoted context omitted.

[flagged]

I would also assume the same for non-Chinese as well

I would assume at this point that any SaaS product, LLM or not, where the data doesn't reside on your servers on your premises (or your own colo) is training on the entire corpus of your data, whether they'll admit to it or not.

Re: Kimi K3: Open Frontier Intelligence

#579
post #207
post #50

Earlier quoted context omitted.

> its overall intelligence ranks second only to Claude Fable 5 and GPT-5.6 Sol Pretty sure ranking “second” to two others means ranking third.

While you are technically correct, in English it’s perfectly fine to say it this way as well. “Second only” here has meaning “next after”, not “number two”.

Yes. "Second to" takes a set as an argument in English. Even the empty set works!

England is second to none.

Re: Kimi K3: Open Frontier Intelligence

#580
post #330

Just in case you were thinking of signing up directly with Moonshot to use the service, they appear to train even on API use: > We may use Content to provide, maintain, develop, support, and improve the Services, comply with applicable law, enforce our terms and policies, and keep the Services safe and secure. Customer who requires restrictions on the use of Customer Content for training or improving Moonshot AI mode…

[flagged]

The reason why the models are getting better is training on users conversations.
Post reply on HN