Live data from Hacker News

Kimi K3: Open Frontier Intelligence

kimi.com

761–770 of 1001 posts

Re: Kimi K3: Open Frontier Intelligence

#761

Earlier quoted context omitted.

Hey Simon, I noticed one thing all LLMs are currently pretty bad at and maybe we could create a benchmark from it. Let an LLM play the role of a dungeon master and tell it to strictly stay in the script/story and only allow realistic player actions. You will notice that they are easily brought off track. E.g. - Tell the LLM that you as a player noticed a strange glow in an NPCs eyes -> the NPC becomes an enemy. - In…

This is a known and solved problem. Such a test is pointless for a general-purpose model, because like most people you're using multiturn chats in a naive way, fighting the default finetuning that is done intentionally. 1. You're sending your in-character inputs to an instruction-tuned model under the user role, in a multiturn chat. It's biased to treat these inputs as instructions and this behavior will show itself…

This is like saying coding is a known and solved problem and benchmarks for coding is pointless for a general model. And that using a model for coding is an abuse of the multiturn chats AIs are tuned for.

1. The user should be able to prompt the AI to act differently from its default behavior. A human assistant is capable of role playing without always sounding like an assistant.

2. If the user asks the AI to follow the script and not allow unrealistic things to happen it should push back. The user is not always absolutely correct.

Re: Kimi K3: Open Frontier Intelligence

#762
post #111

Pelican: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... - rendered via the OpenRouter API: https://openrouter.ai/moonshotai/kimi-k3 95 input, 16,658 output = 25 cents! https://www.llm-prices.com/#it=95&ot=16658&ic=3&oc=15 (13,241 of those were reasoning tokens.) I think that's the most expensive pelican I've rendered through a Chinese model so far.

They should include peican on bike on the release page or model card, alongside those barcharts with Fable

At this point it is useless benchmark because it can just be embedded.

Re: Kimi K3: Open Frontier Intelligence

#763

Earlier quoted context omitted.

so it is ~ same price as openai, same score, but somehow it is mediocre? edit: not to mention being an open model that you can host yourself

Theres no way you will be able to host it yourself. It’s way too large.

Not for any you. But there exists a you for which you can host it yourself.

Re: Kimi K3: Open Frontier Intelligence

#764

I took advantage of their "Token Cup" for the world cup and won 530,000 credits. I believe at the time they said it had to be used in the desktop app, which I have installed. Nowhere can I find any sort of balance or evidence of the 530k other than the Token Cup page itself that say that is what I was given. Their web chat has almost no settings of customization. Everything they present just comes off as amateurish t…

Yea i got about 900k - they have since added a "Gift usage" section under settings/ My Quota

The kimi.com interface also seems to indicate they can be used there (the badge to say its using gift quota is there for me).

However under Usage Details/Gift Quota it seems to indicate that it is consuming it via kimi code and sure enough my usage is reflected there from kimi cli. Odd and a tad vague

Re: Kimi K3: Open Frontier Intelligence

#765
post #584

Earlier quoted context omitted.

Frontier labs release frontier models to the public only if there is market pressure to do so. Anthropic is not even hiding that they have been using Mythos internally for months now. I wouldn’t be surprised if OpenAI (so much for “open”) is using GPT-6 internally already. It appears that peasants like us are not going to get access to frontier AI anymore at any price.

Anthropic had Mythos-Preview for many months internally, but from available sources was an active work in progress, and it seems they started releasing it via Project Glasswing to partners before the final checkpoint was available.

Maybe or maybe not. Anthropic made it a marketting thing.

Re: Kimi K3: Open Frontier Intelligence

#766
post #518

So Chinese labs are driving essentially towards commodotized intelligence. Even if its a few months behind the US. Is this a classic 'commoditize my compliment' situation? They want to sell the hardware and infrastructure behind AI and make the software part not the value driver / moat? I can see it. But also even two Chinese labs sinking 100s of millions USD into training isn't exactly commoditization. It's still a…

If there was some grand strategy for all Chinese labs, surely it'd have leaked by now. I think its more likely that: - Companies can still make money from commodities - Chinese labs only have 5-10% the valuation of OpenAI/Anthropic, so massive monopoly profits aren't necessary. Profit expectations for tech companies in China are really low in general, complete opposite of the US. - Open weighting is a great way to ge…

-Chinese labs only have 5-10% the valuation of OpenAI/Anthropic

Another reason could be that they want to close the gap between their valuations and OpenAI/Anthropic.

Either Chinese labs are worth more or OpenAI/Anthropic are worth less. One of those is true

Re: Kimi K3: Open Frontier Intelligence

#767
post #111

Pelican: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... - rendered via the OpenRouter API: https://openrouter.ai/moonshotai/kimi-k3 95 input, 16,658 output = 25 cents! https://www.llm-prices.com/#it=95&ot=16658&ic=3&oc=15 (13,241 of those were reasoning tokens.) I think that's the most expensive pelican I've rendered through a Chinese model so far.

> I think that's the most expensive pelican I've rendered through a Chinese model so far.

quite insane that it costs as much as 5.6 Terra [1], and twice the European counterpart (albeit dated for today's standards?) [2].

to be fair, the pelicans from Terra were quite weird all things considered. also, given the limited TPS from the first-party, it has to be pushing the limits of inference capabilities.

[1] https://openrouter.ai/openai/gpt-5.6-terra

[2] https://openrouter.ai/mistralai/mistral-medium-3-5

Re: Kimi K3: Open Frontier Intelligence

#768

Earlier quoted context omitted.

In regards to your post and the 16k reasoning output. Try setting reasoning levels yourself manually. We see in the benchmarks that one of the graphs shows low, mid, max, so its clearly there. I had the same issue with GLM 5.2 only offering high/max. By playing around with openai compatible protocol, and setting the reasoning level from none, low ... high, xhigh and testing a flawed logic test. It was easy to see tha…

Official API doc says only max effort level is currently supported. https://platform.kimi.ai/docs/guide/kimi-k3-quickstart#think...

[deleted]

Re: Kimi K3: Open Frontier Intelligence

#769
post #111

Pelican: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... - rendered via the OpenRouter API: https://openrouter.ai/moonshotai/kimi-k3 95 input, 16,658 output = 25 cents! https://www.llm-prices.com/#it=95&ot=16658&ic=3&oc=15 (13,241 of those were reasoning tokens.) I think that's the most expensive pelican I've rendered through a Chinese model so far.

Hey Simon, I noticed one thing all LLMs are currently pretty bad at and maybe we could create a benchmark from it. Let an LLM play the role of a dungeon master and tell it to strictly stay in the script/story and only allow realistic player actions. You will notice that they are easily brought off track. E.g. - Tell the LLM that you as a player noticed a strange glow in an NPCs eyes -> the NPC becomes an enemy. - In…

This is a fascinating example of perhaps why a move towards a "world model" or some better form of representation can be helpful.

Re: Kimi K3: Open Frontier Intelligence

#770
post #111

Pelican: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... - rendered via the OpenRouter API: https://openrouter.ai/moonshotai/kimi-k3 95 input, 16,658 output = 25 cents! https://www.llm-prices.com/#it=95&ot=16658&ic=3&oc=15 (13,241 of those were reasoning tokens.) I think that's the most expensive pelican I've rendered through a Chinese model so far.

> I think that's the most expensive pelican I've rendered through a Chinese model so far. quite insane that it costs as much as 5.6 Terra [1], and twice the European counterpart (albeit dated for today's standards?) [2]. to be fair, the pelicans from Terra were quite weird all things considered. also, given the limited TPS from the first-party, it has to be pushing the limits of inference capabilities. [1] https://op…

I don’t think a 128B model will be that competitive with a 2.8T model. If anything, one should wonder why Mistral is so expensive in the current day.
Post reply on HN