Live data from Hacker News

The Kimi K3 Moment

stephen.bochinski.dev

611–620 of 644 posts

Re: The Kimi K3 Moment

#611

Earlier quoted context omitted.

my first prompt to any Kimi model was K3 via Pi, some version of "hi kimi!!" and the response was telling me "I'm actually Claude." this is not hard to repro, just use a system prompt that doesn't mention the model name. that said, if they bootstrapped with opus 4.6 convo sft data they had sitting around... so what?

The main story is what isn't being talked about. Chinese labs exfiltrated trillions of tokens of high-quality output from Anthropic and OpenAI, through proxies and heavily discounted token resellers, which they distilled and used for training data for their own models. Instead of spending 12-18 months building their own robust harnesses and painstakingly creating quality training data (which is what Anthropic and Ope…

> We witnessed the most extensive industrial espionage campaign, probably ever

This is the funniest way of saying “going to a company’s website” I have seen in my entire life

Re: The Kimi K3 Moment

#612

Earlier quoted context omitted.

First off, I don't doubt that China engages in large-scale industrial espionage, or at least used to. Nowadays they have plenty of talented and highly educated engineering talent of their own. Second, I now actually read the article. It describes plenty of questionable and problematic things but also contradicts your claims explicitly. The essential point of the article is about making money by selling access to Clau…

> Nothing in the article suggests or supports the idea of large-scale coordinated "distillation attacks" Did you read the correct article? This is covered in the first paragraph: https://www.anthropic.com/news/detecting-and-preventing-dist... It directly addresses large-scale, coordinated 'distillation attacks' orchestrated by Chinese labs, which Anthropic accuses of exfiltrating tens of millions of exchanges. The re…

I enjoy your detailed breakdown of how exactly they pay anthropic for access to their models but I am unclear about how “I think they’re bad” means that they didn’t pay for it

Can you elaborate on “it doesn’t count as a sale if I don’t like them even if I accept the money and give them what they paid for”? Like if you work at a gas station and you realize that the guy that just bought a hot dog bullied you in middle school do you call the cops?

Re: The Kimi K3 Moment

#613
post #115

Earlier quoted context omitted.

Correct: can't opt out of training. This is well documented. "Can't use for commercial purposes" - incorrect AFAICT. In what sense do you mean this? The open weight MIT version obviously allows for commercial use, but I don't think that's what you're referring to, because training data is irrelevant on the open weight version. Pretty sure the API allows commercial use too. Maybe the free version doesn't? But who care…

> Service Misuse. You acknowledge that without the written consent of us and/or the relevant rights holders, (i)you have no authority to use Kimi and the content generated by Kimi in any commercial manner; (ii)you may not use our Services to develop products or services that compete with us. https://www.kimi.com/user/agreement/modelUse

That’s for the consumer app / chatbot. For the api the terms are different: https://platform.kimi.ai/docs/agreement/modeluse

Re: The Kimi K3 Moment

#614

Earlier quoted context omitted.

API distillation doesn't have to explain all of K3's capabilities for it to have happened. Kimi K3 reproducibly identifies itself as Claude: https://x.com/denisewu/status/2077984660211269870 This behavior is exactly what you'd expect from a model distilled from Claude. There's a detailed analysis of K3's ambiguous identity here: https://github.com/rgreenblatt/which_claude_is_k3/blob/main/... This analysis observed K3…

But then again, the identity could also have slipped into the model from other sources during pretraining. The internet is full of "I am Claude": https://grep.app/search?q=i+am+claude and variants https://grep.app/search?q=i%27m+claude Either way, there's probably no significant portion of Mythos/Fable or Sol in there as OP has stated.

When prefaced with "I am Claude", Kimi K3 prefers to generate API-specific Anthropic model identifiers, unlike other models Qwen, GPT, or even Claude itself. These exact identifiers appear in Claude API metadata, and are stripped out of Claude web chats.

While other models produce human-readable names like "Opus 4.5" or "Sonnet 4", Kimi K3 produces exact API model identifier like "claude-opus-4-5-20251101" or "claude-sonnet-4-20250514".

Which is extremely unusual. Web chats only contain the human-readable model name. Other models don't do this. So where did K3 get this data?

We can conclude, with high confidence, that:

1) K3 was trained on raw Claude API calls/metadata.

2) Claude API metadata was trained on in additional to standard web data.

Re: The Kimi K3 Moment

#615

Earlier quoted context omitted.

Sure, but then Qwen should leak that too, and it doesn't. K3 calls itself Claude 7 out of 48 times, Qwen does it 0 out of 48, and the only other model to identify itself as Claude is DeepSeek. and DeepSeek is alleged to also distill from Claude data anyway. So this isn't something every model absorbed from the same web text. And you skipped over the strongest datapoint that K3 is distilled: K3 reproduces Claude's pub…

> Qwen should leak that too, and it doesn't. FWIW I had Qwen identifying itself as "a language model made by Google" in one conversation, although I could not reproduce this reliably.

here's a chart of asking various models to self-identify: https://github.com/rgreenblatt/which_claude_is_k3/blob/main/...

Qwen is remarkably consistent, correctly identifying itself 100% of the time. Kimi K3 performs poorly at this test, it correctly self-identifies ~80% of the time; sometimes calling itself Claude and rarely ChatGPT.

Re: The Kimi K3 Moment

#616
post #442

Earlier quoted context omitted.

> If you ever listen to Russian propaganda, there's a similar theme: every big idea, everything good, all of it was definitely first developed in Russia - only Russians could ever have thought of it. Of course, Russia isn't actually a world leader in any of those things, or able to execute on them. When I was a kid watching Star Trek VI, I was confused by the line "You've not experienced Shakespeare until you've read…

I mean, lots of things have been developed in European countries, and said countries either never commercialized the technology well enough, or lost leadership. I'm sure that prior to WW2, Brits would've treated the idea that they could be out-engineered by the Germans with ridicule.

I'm not sure what you're arguing here.

Upthread, the point about it not mattering what China stole but what they can do? That "Made in China" has gone from a sign of low quality to being the default as it became the factory of the world? Yes, China's winning and the sooner the rest of us wake up and smell the tea the better. They learned lessons from how my ancestors were able to push them around despite coming from a small, damp, sheep-filled rock in the Atlantic, and don't want that to happen again; for most of "the west", being on the receiving end of such humiliation is a historical footnote if it's in our own history books (or living history) at all*, and most of the exceptions are former-Soviet-bloc/Warsaw Pact.

But specifically to the point about the USSR claiming all (good) things for itself, as parodied in this manner by Star Trek? AFAICT, that was just plain soviet propaganda.

* What happened in Africa and India/Pakistan, however, is recent enough to still be in living memory. Israel likewise, though now this is on the edge of living memory for the things which led to its reincarnation.

And while the Irish definitely still retell history lessons about the British and the Famine, I'm not at all sure if the stuff in NI in my lifetime counts as "humiliation" given I'm British and our newspapers just said "terrorist" about everything that came from there.

This is only relevant to the point that most of us are deeply oblivious to how bad things can get when someone else is the boss and we don't get to vote in their elections.

Re: The Kimi K3 Moment

#617

Earlier quoted context omitted.

But then again, the identity could also have slipped into the model from other sources during pretraining. The internet is full of "I am Claude": https://grep.app/search?q=i+am+claude and variants https://grep.app/search?q=i%27m+claude Either way, there's probably no significant portion of Mythos/Fable or Sol in there as OP has stated.

When prefaced with "I am Claude", Kimi K3 prefers to generate API-specific Anthropic model identifiers, unlike other models Qwen, GPT, or even Claude itself. These exact identifiers appear in Claude API metadata, and are stripped out of Claude web chats. While other models produce human-readable names like "Opus 4.5" or "Sonnet 4", Kimi K3 produces exact API model identifier like "claude-opus-4-5-20251101" or "claude…

The exact model identifiers appear extremely frequently in code on GitHub.

https://grep.app/search?q=claude-opus-4-5-20251101

https://grep.app/search?q=claude-sonnet-4-20250514

They also appear elsewhere on the internet:

https://trends.google.com/trends/explore?q=claude-opus-4-5-2...

Re: The Kimi K3 Moment

#618
Does this article really compare a single well defined LLM model "Kimi K3" vs. a family of LLM models including Haiku x.x, Sonnet x.x, Opus x.x, Fable x.x without actually revealing what Claude-family model was being compared?

(Fable has been restricted somewhat, but the article uses Fable pricing as a comparison point, so it worth including it in the list of possible Claude family LLMs)

It's hard to know what to take away from the post with this ambiguity.

It's also worth noting the range of pricing for the Claude family ranges from $1/$5 a million in/out to $10/$50 a million in/out, so the ambiguity of which particular model the comparison is against spans a 10x range of model fees.

Re: The Kimi K3 Moment

#619

Pricing is actually far cheaper than that. There's two tiers of pricing: Chinese and US. If you sign up with non-Chinese phone number, you're bucketed into US, you get US prices, can pay only in USD and with American credit card network. Chinese prices are about 9x cheaper than the US prices, which are already far cheaper than Claude or other American provider. If you can somehow get hold of a Chinese phone number, k…

Most of this hand-wringing on price will go away. My assumption is that Anthropic, OpenAI, Kimi, etc all have a similar cost structure when serving models. The same size model roughly generates the same GPU usage whether you’re American or Chinese. I’d also guess that the model sizes across all SOTA models is similar, we just only see data for open models. The difference is most likely that American companies simply…

It’s about the cost of energy. And China has cheap, clean energy.

The calculation is about tokens/gigawatt and $/gigawatt.

Re: The Kimi K3 Moment

#620
post #616

Earlier quoted context omitted.

I mean, lots of things have been developed in European countries, and said countries either never commercialized the technology well enough, or lost leadership. I'm sure that prior to WW2, Brits would've treated the idea that they could be out-engineered by the Germans with ridicule.

I'm not sure what you're arguing here. Upthread, the point about it not mattering what China stole but what they can do? That "Made in China" has gone from a sign of low quality to being the default as it became the factory of the world? Yes, China's winning and the sooner the rest of us wake up and smell the tea the better. They learned lessons from how my ancestors were able to push them around despite coming from…

well speaking as a chinese, we only want our relics back.....nothing else from UK
Post reply on HN