The Kimi K3 Moment
621–630 of 644 posts
Re: The Kimi K3 Moment
#622Earlier quoted context omitted.
assume you are a "second class lab" and you are in fact making progress by distilling the results of the frontier labs' efforts. what is the end game for this strategy? if the frontier labs shut down, or stop releasing to the public, and there's noting left to distill, how will you progress?
Distillation from a teacher model solves the self-start problem, that is, building a model to the point where it reason coherently. Without distillation, solving self-start is incredibly difficult since it requires millions of high quality training samples. Creating that kind of dataset takes an enormous amount of effort. Once a model becomes competent enough to perform complex reasoning, a teacher model is no longer…
Can you imagine the amount of effort it takes to write 15 trillion tokens worth of art, literature, source code, textbooks, scientific papers, news articles, etc.? No wonder Anthropic just scooped it up and took it for free!
Yet you don’t seem bothered by this. I wonder why.
Re: The Kimi K3 Moment
#623Earlier quoted context omitted.
> Nothing in the article suggests or supports the idea of large-scale coordinated "distillation attacks" Did you read the correct article? This is covered in the first paragraph: https://www.anthropic.com/news/detecting-and-preventing-dist... It directly addresses large-scale, coordinated 'distillation attacks' orchestrated by Chinese labs, which Anthropic accuses of exfiltrating tens of millions of exchanges. The re…
I enjoy your detailed breakdown of how exactly they pay anthropic for access to their models but I am unclear about how “I think they’re bad” means that they didn’t pay for it Can you elaborate on “it doesn’t count as a sale if I don’t like them even if I accept the money and give them what they paid for”? Like if you work at a gas station and you realize that the guy that just bought a hot dog bullied you in middle…
Re: The Kimi K3 Moment
#624Earlier quoted context omitted.
> Kimi K3 reproducibly identifies itself as Claude It could also be have been trained from collected response datasets. Claude got caught several time responding it was ChatGPT or even Deepseek and I don't think Anthropic has been distealling DeepSeek. > This behavior is exactly what you'd expect from a model distilled from Claude. The opposite actually. If they wanted to distill Claude without getting caught they co…
> they could just use a regex to change Claude to Kimi in their distillation pipeline! Jean-Kimi Van Damme would like to have a word with you.
Re: The Kimi K3 Moment
#625Earlier quoted context omitted.
When prefaced with "I am Claude", Kimi K3 prefers to generate API-specific Anthropic model identifiers, unlike other models Qwen, GPT, or even Claude itself. These exact identifiers appear in Claude API metadata, and are stripped out of Claude web chats. While other models produce human-readable names like "Opus 4.5" or "Sonnet 4", Kimi K3 produces exact API model identifier like "claude-opus-4-5-20251101" or "claude…
The exact model identifiers appear extremely frequently in code on GitHub. https://grep.app/search?q=claude-opus-4-5-20251101 https://grep.app/search?q=claude-sonnet-4-20250514 They also appear elsewhere on the internet: https://trends.google.com/trends/explore?q=claude-opus-4-5-2...
~3000 instances over the entirety of GitHub is not "extremely frequent" at all when you consider massive scale of the pretraining corpus. These models are trained on all text on the accessible internet plus millions of books, on trillions of words overall. ~3000 instances isn't even a rounding error.
Also, Google Trends is the wrong tool for this case. The link you replied with measures the number of Google searches for a specific query, which is completely irrelevant to how often a string actually appears on the web.
I looked at those links. They show random GitHub code samples which contain this model identifier. Even if K3 did train on those GitHub code samples, those are random code strings and K3 isn’t just reciting random pieces of code. When prefilled with "I am Claude", K3 answers with Anthropic's exact backend API identifiers, instead of the human conversational-style name "Opus 4.5".
If this was a result of scraping the open web and GitHub, other frontier models trained on GitHub data (like Qwen, Claude, GPT) would show a similar API model name when prompted. But they don't. Only K3 does this behavior, and it prefers to answer with the exact API tags like "claude-opus-4-5-20251101" and "claude-sonnet-4-20250514".
A LLM doesn't suddenly output a rare API model identifier when prompted with "I am Claude", just because it saw it in a .py file. No, it only does this because that identifier must have appeared rather frequently in the training data. Which is what happens when you distill Claude models, and don't properly clean your distillation data.
Plus, this isn't just spurious speculation that Moonshot is distilling Claude. Anthropic caught Moonshot as running "industrial-scale" distillation campaign from Claude: https://www.anthropic.com/news/detecting-and-preventing-dist...
To quote:
Moonshot (Kimi models) employed hundreds of fraudulent accounts spanning multiple access pathways. Varied account types made the campaign harder to detect as a coordinated operation. We attributed the campaign through request metadata, which matched the public profiles of senior Moonshot staff. In a later phase, Moonshot used a more targeted approach, attempting to extract and reconstruct Claude’s reasoning traces.
The operation targeted:
* Agentic reasoning and tool use
* Coding and data analysis
* Computer-use agent development
* Computer vision
Re: The Kimi K3 Moment
#626Earlier quoted context omitted.
Sure, but then Qwen should leak that too, and it doesn't. K3 calls itself Claude 7 out of 48 times, Qwen does it 0 out of 48, and the only other model to identify itself as Claude is DeepSeek. and DeepSeek is alleged to also distill from Claude data anyway. So this isn't something every model absorbed from the same web text. And you skipped over the strongest datapoint that K3 is distilled: K3 reproduces Claude's pub…
Did you know claude models identify as qwen or deepseek when asked in chinese?
For those asserting that Kimi identifying as another model is evidence of distillation, is this evidence that Claude was distilled from Chinese models?
Re: The Kimi K3 Moment
#627Earlier quoted context omitted.
How is the efficient market hypothesis applicable here?
Your point that “people will be gobsmacked” suggests you think you know something that the broad market does not. In a loose way, that implies the market has enough information to anticipate this but is not pricing assets accordingly. Asked another way, why not just short American assets if you are convinced of your hypothesis? Why live here, assuming you do?
Re: The Kimi K3 Moment
#628Earlier quoted context omitted.
How is the efficient market hypothesis applicable here?
The efficient market hypothesis, read loosely, says that capitalism is the best system. Yet here it is being thoroughly pwned by a series of 5-year plans.
The modern Chinese system embraces Econ 101 "capitalism" in the sense that it relies on markets for price discovery. Today, the "5 year plans" are more like what we would call "industrial policy" in the west. That's different than pure capitalism, but so is the American system. Alexander Hamilton and Abraham Lincoln both advocated strong federal intervention in the economy in service of industrial policy: https://emergingamerica.org/blog/alexander-hamilton-founder-... ("Hamilton’s plan involved the following aspects: 1) the creation of a Federally backed source of credit, the First Bank of the United States; 2) a system of tariffs, bounties, and other financial incentives or penalties to support the “essential” sectors of the U.S. economy; and 3) Federal support for developing manufactures by helping fund physical infrastructure ( transportation, in particular), regulating quality standards, and creating an institution to promote 'the prosecution and introduction of useful discoveries, inventions and improvements.'").
Re: The Kimi K3 Moment
#629Earlier quoted context omitted.
First off, I don't doubt that China engages in large-scale industrial espionage, or at least used to. Nowadays they have plenty of talented and highly educated engineering talent of their own. Second, I now actually read the article. It describes plenty of questionable and problematic things but also contradicts your claims explicitly. The essential point of the article is about making money by selling access to Clau…
> Nothing in the article suggests or supports the idea of large-scale coordinated "distillation attacks" Did you read the correct article? This is covered in the first paragraph: https://www.anthropic.com/news/detecting-and-preventing-dist... It directly addresses large-scale, coordinated 'distillation attacks' orchestrated by Chinese labs, which Anthropic accuses of exfiltrating tens of millions of exchanges. The re…
https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens...
I know that Anthropic claims that "large-scale coordinated "distillation attacks"" exist. They are not a disinterested party here and they don't provide evidence either.
Also you're just moving goalposts around. The article you linked also cites Anthropics claim but then paints a much much mor nuanced picture.
Re: The Kimi K3 Moment
#630Earlier quoted context omitted.
This is essentially the point of the visa, it feels wrong especially as YC drops standards and increases cohort sizes, but the same power laws that keep them winning also apply here in maximizing economic value of each O-1 approval Basically of all visas O-1 is virtually guaranteed to have highly positive economic value
It's not the point of the visa, the O-1 is supposed to be for people of extraordinary ability, eg Nobel Prize winners. It's used for software engineers.