Live data from Hacker News

The Kimi K3 Moment

stephen.bochinski.dev

371–380 of 644 posts

Re: The Kimi K3 Moment

#371

Earlier quoted context omitted.

I strongly agree with the premise that distillation is not an “attack”. But that said: K3 is not a distilled version of Fable or Sol. Fable has been barely available and Sol was just released! Moreover, K3 is superior to both models in some domains, according to user scoring on the Arena. API distillation can’t give you these results anyway. All it is useful for is bootstrapping RL in new domains to get past the “col…

API distillation doesn't have to explain all of K3's capabilities for it to have happened. Kimi K3 reproducibly identifies itself as Claude: https://x.com/denisewu/status/2077984660211269870 This behavior is exactly what you'd expect from a model distilled from Claude. There's a detailed analysis of K3's ambiguous identity here: https://github.com/rgreenblatt/which_claude_is_k3/blob/main/... This analysis observed K3…

But then again, the identity could also have slipped into the model from other sources during pretraining. The internet is full of "I am Claude": https://grep.app/search?q=i+am+claude and variants https://grep.app/search?q=i%27m+claude

Either way, there's probably no significant portion of Mythos/Fable or Sol in there as OP has stated.

Re: The Kimi K3 Moment

#372

Earlier quoted context omitted.

Your main source is Ryan Greenblatt who is a regular recipient of community notes and has no corroboration for the 15% statistic other than his assertion. The other tweet (Sauers_) is also community noted as engagement farming with a false system prompt, so forgive me for being skeptical.

Including the three sources above, multiple others have reported that K3 self-identifies as Claude. "I'm actually Claude - not Kimi". https://x.com/PimDeWitte/status/2077884701470040083 I regret to inform you that it is, in fact, real and from their own website - you don’t even need to try hard to reproduce it. https://x.com/PimDeWitte/status/2078105292965912690 lmao this is so funny, if you ask Kimi K3 for something…

Moonshot AI should have made it identify as Mythos as a practical joke to make US go crazy trying to figure out how they got access to it.

Re: The Kimi K3 Moment

#373

Earlier quoted context omitted.

Including the three sources above, multiple others have reported that K3 self-identifies as Claude. "I'm actually Claude - not Kimi". https://x.com/PimDeWitte/status/2077884701470040083 I regret to inform you that it is, in fact, real and from their own website - you don’t even need to try hard to reproduce it. https://x.com/PimDeWitte/status/2078105292965912690 lmao this is so funny, if you ask Kimi K3 for something…

my first prompt to any Kimi model was K3 via Pi, some version of "hi kimi!!" and the response was telling me "I'm actually Claude." this is not hard to repro, just use a system prompt that doesn't mention the model name. that said, if they bootstrapped with opus 4.6 convo sft data they had sitting around... so what?

The main story is what isn't being talked about. Chinese labs exfiltrated trillions of tokens of high-quality output from Anthropic and OpenAI, through proxies and heavily discounted token resellers, which they distilled and used for training data for their own models.

Instead of spending 12-18 months building their own robust harnesses and painstakingly creating quality training data (which is what Anthropic and OpenAI did), they distilled Anthropic's models to bypass the hardest parts of development. Chinese labs compressed 18 months of intensive research and development into just 6 months, and are now head-to-head with their American counterparts.

Anthropic tried to complain about this unauthorized "token theft", but they burned too much public goodwill with BS safety restrictions and users don't care. The US government is too busy fighting a war to help. Chinese labs are offering highly capable, cheap, open-weight models; exactly what users want. The community is happy to overlook any questionable methods Chinese labs used to build them.

The cope is incredible. There's people in this thread in denial that Moonshot AI is trained on exfiltrated Anthropic's model output, even when shown substantial evidence this has been happening since Kimi 2.X

Chinese labs were even paying an absurd $0.01 per Opus tool call trace, to get the quantity of training data needed.

Kimi K3 has reached the point of RSI, and no longer needs synthetic data generated by Anthropic/OpenAI models. K3 is now capable enough to generate, iterate, and improve its own training data recursively. The data exfiltration is complete.

We witnessed the most extensive industrial espionage campaign, probably ever, and nobody in the industry cares at all that it happened.

Re: The Kimi K3 Moment

#374

Earlier quoted context omitted.

my first prompt to any Kimi model was K3 via Pi, some version of "hi kimi!!" and the response was telling me "I'm actually Claude." this is not hard to repro, just use a system prompt that doesn't mention the model name. that said, if they bootstrapped with opus 4.6 convo sft data they had sitting around... so what?

The main story is what isn't being talked about. Chinese labs exfiltrated trillions of tokens of high-quality output from Anthropic and OpenAI, through proxies and heavily discounted token resellers, which they distilled and used for training data for their own models. Instead of spending 12-18 months building their own robust harnesses and painstakingly creating quality training data (which is what Anthropic and Ope…

assuming the k3 model weights do indeed get published, if your model of the world is "achieving RSI is beneficial and K3 has done so," this feels structurally different from ordinary industrial espionage, because the knowledge has enriched the commons

more like silk than capacitors

if, again, your model is that RSI will be beneficial, why wouldn't making it available to all unlock more benefit globally than not doing that

Re: The Kimi K3 Moment

#375

Earlier quoted context omitted.

I strongly agree with the premise that distillation is not an “attack”. But that said: K3 is not a distilled version of Fable or Sol. Fable has been barely available and Sol was just released! Moreover, K3 is superior to both models in some domains, according to user scoring on the Arena. API distillation can’t give you these results anyway. All it is useful for is bootstrapping RL in new domains to get past the “col…

API distillation doesn't have to explain all of K3's capabilities for it to have happened. Kimi K3 reproducibly identifies itself as Claude: https://x.com/denisewu/status/2077984660211269870 This behavior is exactly what you'd expect from a model distilled from Claude. There's a detailed analysis of K3's ambiguous identity here: https://github.com/rgreenblatt/which_claude_is_k3/blob/main/... This analysis observed K3…

fwiw, Gemini 3.5 has identified itself to me as an OpenAI product on multiple occasions.

Re: The Kimi K3 Moment

#376

This was always where this was heading, but we got here much faster than expected. Once western governments declare it to be a "national security" risk for citizens to have access to open-weight frontier models, and once they classify using these models as acts of terrorism, what will that world be like? Will using Kimi K3 come to be like how napster was in the olden days? Everybody knew it was technically illegal, b…

The disappearance of high ram Mac studio rigs is probably just a coincidence, right? :/

More like the pricing was too turbulent. Apple doesn’t really do rapidly changing prices and there isn’t any stable price for that much ram.

Re: The Kimi K3 Moment

#377

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

It is unfair, they stole the dataset that we stole.

Re: The Kimi K3 Moment

#378

Earlier quoted context omitted.

Not sure how the economics work for the Chinese models, but DeepSeek did the same task for a dime.

In my opinion, for the vast majority of use cases, DeepSeek is still the most cost-effective model by a mile. $10 feels like it lasts forever.

Maybe I’m not pushing them hard enough but I use Claude opus at work and deepseek v4 flash at home and they both seem about as capable. While deepseek is borderline free.

Re: The Kimi K3 Moment

#379
post #116

Earlier quoted context omitted.

At this point, the United States will lose that battle most Countries in the world are going end up using electronics from Asia, that ship has sailed Japan, China, Korea, Singapore, Taiwan, Vietnam, dominate that area, China already dominates EVs, Drones and many other electronic devices, and with the way Donald Trump has picked fights, Europe, Canada, Australia, New Zealand, Mexico and many others are looking for ot…

[flagged]

We've banned this account for using HN for ideological battle and ignoring our requests to stop.

Re: The Kimi K3 Moment

#380

I think it's the opposite. Kimi K3 has 2.8 trillion parameters. We don't know the number of parameters of ChatGPT 5.6 or Opus 4.8, but it's probably in the same region. Fable/Mythos are rumored to be around 10 trillion. So, K3 is directly comparable with ChatGPT 5.6 and Opus 4.8, and the price is not so much lower: K3: $3/$15 per 1 Mtok input/output ChatGPT 5.6 Sol: $5/$30 Opus 4.8: $5/$25 This is not a watershed mom…

I’d also note that running a 2.8 trillion parameter model at scale efficiently is not simple. I would expect when open weights land getting it running fast, efficient, and at full capability will require sufficient resources it’ll be expensive outside of Chinese hosting. Which I think almost no western corporation would use for any internal work. You have to anticipate your use won’t just go towards training but will be actively mined for IP, trade secrets, MNPI, etc, or anything of use to the Chinese government or Chinese companies. I don’t say this to crap on the Chinese - but this is the playbook for the last 30 years.

That said I fully intend to use deepseek hosting for operational agents that are making decisions about non sensitive material. The economics are astounding.

Kimi? The economics aren’t that amazing to merit switching from 5.6. I expect fable will rapidly reappear in subscriptions. Competition is good.

Post reply on HN