Earlier quoted context omitted.
The desire to accuse China of just copying is like 20 years out of date. It’s been wrong since some people on HN were in diapers. People are going to be gobsmacked when, in our lifetime, China becomes a world power comparable to the U.S. Probably still poorer per capita, but at Spain/Italy levels, not third world country levels. And they’ll be shocked at the implications of that on the world economy, migration patter…
So the efficient market hypothesis is wrong?
The Kimi K3 Moment
381–390 of 644 posts
Re: The Kimi K3 Moment
#382Earlier quoted context omitted.
my first prompt to any Kimi model was K3 via Pi, some version of "hi kimi!!" and the response was telling me "I'm actually Claude." this is not hard to repro, just use a system prompt that doesn't mention the model name. that said, if they bootstrapped with opus 4.6 convo sft data they had sitting around... so what?
The main story is what isn't being talked about. Chinese labs exfiltrated trillions of tokens of high-quality output from Anthropic and OpenAI, through proxies and heavily discounted token resellers, which they distilled and used for training data for their own models. Instead of spending 12-18 months building their own robust harnesses and painstakingly creating quality training data (which is what Anthropic and Ope…
Hard to feel sorry for companies that created their empires by ignoring copyright themselves.
Also, 'most extensive industrial espionage campaign, probably ever' is absolute nonsense. They did not need to infiltrate the companies for this nor are you accusing them of stealing any trade secrets. This is only about whether they looked at their competitors' products from the outside (in the form of conversation tokens) and used it to improve their own product (by training). Hardly the crime of the century.
Re: The Kimi K3 Moment
#383Earlier quoted context omitted.
my first prompt to any Kimi model was K3 via Pi, some version of "hi kimi!!" and the response was telling me "I'm actually Claude." this is not hard to repro, just use a system prompt that doesn't mention the model name. that said, if they bootstrapped with opus 4.6 convo sft data they had sitting around... so what?
The main story is what isn't being talked about. Chinese labs exfiltrated trillions of tokens of high-quality output from Anthropic and OpenAI, through proxies and heavily discounted token resellers, which they distilled and used for training data for their own models. Instead of spending 12-18 months building their own robust harnesses and painstakingly creating quality training data (which is what Anthropic and Ope…
Re: The Kimi K3 Moment
#384Earlier quoted context omitted.
my first prompt to any Kimi model was K3 via Pi, some version of "hi kimi!!" and the response was telling me "I'm actually Claude." this is not hard to repro, just use a system prompt that doesn't mention the model name. that said, if they bootstrapped with opus 4.6 convo sft data they had sitting around... so what?
The main story is what isn't being talked about. Chinese labs exfiltrated trillions of tokens of high-quality output from Anthropic and OpenAI, through proxies and heavily discounted token resellers, which they distilled and used for training data for their own models. Instead of spending 12-18 months building their own robust harnesses and painstakingly creating quality training data (which is what Anthropic and Ope…
Oh, no. I wouldn’t say that. If that happened, I definitely care: I’m positively delighted about it.
Re: The Kimi K3 Moment
#385This was always where this was heading, but we got here much faster than expected. Once western governments declare it to be a "national security" risk for citizens to have access to open-weight frontier models, and once they classify using these models as acts of terrorism, what will that world be like? Will using Kimi K3 come to be like how napster was in the olden days? Everybody knew it was technically illegal, b…
If they outlaw open source models that'll just handicap American companies, the rest of the world will be running open source and have an arbitrage against US companies.
Re: The Kimi K3 Moment
#386I think it's the opposite. Kimi K3 has 2.8 trillion parameters. We don't know the number of parameters of ChatGPT 5.6 or Opus 4.8, but it's probably in the same region. Fable/Mythos are rumored to be around 10 trillion. So, K3 is directly comparable with ChatGPT 5.6 and Opus 4.8, and the price is not so much lower: K3: $3/$15 per 1 Mtok input/output ChatGPT 5.6 Sol: $5/$30 Opus 4.8: $5/$25 This is not a watershed mom…
Re: The Kimi K3 Moment
#387Earlier quoted context omitted.
This line made me think by 'normal coding work' the author means doing something they don't understand well enough to be able to distinguish the models' output.
Steve Yegge calls this is the "discernment horizon" - https://steve-yegge.medium.com/the-flat-curve-society-36c8b0...
Although in the month since that most recent post, his other points about open models are undercut by K3.
And at least of data available as of 2026-01, AI compute capacity was doubling every 7 months, so I expect every major country to host AI compute farms, and self-host AI feasibility to majorly increase in the next 2-3 years as well. (Partially undercutting, but not fully disproving, his points.)
And I wish his posts were 4 times less wordy.
Re: The Kimi K3 Moment
#388Earlier quoted context omitted.
I strongly agree with the premise that distillation is not an “attack”. But that said: K3 is not a distilled version of Fable or Sol. Fable has been barely available and Sol was just released! Moreover, K3 is superior to both models in some domains, according to user scoring on the Arena. API distillation can’t give you these results anyway. All it is useful for is bootstrapping RL in new domains to get past the “col…
API distillation doesn't have to explain all of K3's capabilities for it to have happened. Kimi K3 reproducibly identifies itself as Claude: https://x.com/denisewu/status/2077984660211269870 This behavior is exactly what you'd expect from a model distilled from Claude. There's a detailed analysis of K3's ambiguous identity here: https://github.com/rgreenblatt/which_claude_is_k3/blob/main/... This analysis observed K3…
It could also be have been trained from collected response datasets. Claude got caught several time responding it was ChatGPT or even Deepseek and I don't think Anthropic has been distealling DeepSeek.
> This behavior is exactly what you'd expect from a model distilled from Claude.
The opposite actually. If they wanted to distill Claude without getting caught they could just use a regex to change Claude to Kimi in their distillation pipeline!
Re: The Kimi K3 Moment
#389Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…
The desire to accuse China of just copying is like 20 years out of date. It’s been wrong since some people on HN were in diapers. People are going to be gobsmacked when, in our lifetime, China becomes a world power comparable to the U.S. Probably still poorer per capita, but at Spain/Italy levels, not third world country levels. And they’ll be shocked at the implications of that on the world economy, migration patter…
Re: The Kimi K3 Moment
#390GPT 5.6 Sol comes out ahead of Kimi K3 on price/task (but not significantly so). You're probably thinking, "Why use Kimi K3? Isn't an open model supposed to beat the closed one on price?", but you need to consider that the closed models are completely hobbled when trying to do anything security-related. For my use-case, I can't risk getting pwned because I'm using a model that refuses to secure my app while there is…
Honestly, it's the only sane way for the market to move. The big labs are obviously stealing our data. Anthropic in particular clean-rooms everything you feed it, even if you opt out, so that it can train on your IP without getting sued. It's a copyright grey area they're abusing because the law has not kept up.