Live data from Hacker News

The Kimi K3 Moment

stephen.bochinski.dev

351–360 of 644 posts

Re: The Kimi K3 Moment

#351

Earlier quoted context omitted.

It’s so funny to me that Anthropic can make claims like this one with zero evidence provided. DeepSeek and others like Minimax are publishing deep research on Multi-Head Latent Attention and Mixture of Experts, Multi-Token Prediction, novel Sparse Attention approaches, I mean they trained long context models on a fraction of the resources and gave everyone the recipe. Chinese labs might not have the funding of labs l…

There's reproducible evidence of Kimi K3 spontaneously identifying itself as Claude https://x.com/denisewu/status/2077984660211269870 This behavior is exactly what you'd expect from a model distilled from Claude. Someone even took the time to analyze Kimi's ambiguous identity, in great detail: https://github.com/rgreenblatt/which_claude_is_k3/blob/main/... And there's an entire Reddit thread discussing this https://w…

By evidence I mean logs, I mean IP addresses, I mean timestamps. They claim millions of requests, let’s see literally any of them?

I don’t consider a tweet by Denise Wu, who works at Anthropic, to be reproducible evidence.

I don’t consider “Caveat: fully AI-generated research.” To be someone taking time to analyze anything in great detail.

Because two AI models produce vaguely similar front-end styles when generating similar prompts I also do not consider to be of much value?

I think this is what I mean when I say the U.S. has its head in the sand. The Chinese labs are releasing ~60 page research reports with citations and analyses and evidence and Anthropic is throwing up defensive blog posts with zilch. I’ve seen more detail in a tech blog from Uber than anything I’ve seen from Anthropic.

Re: The Kimi K3 Moment

#352

Earlier quoted context omitted.

> This behavior is exactly what you'd expect from a model distilled from Claude. This is not at all what I would expect because it's trivial to change the training data to replace Claude with Kimi. In fact I'd argue it's almost certainly not saying that due to distillation.

I encourage you to review the links before committing to a position. The writeup on K3's anomalous trans-model identity is very comprehensive. K3 reproduces Claude's internal model identifier when prompted, something which the real Claude models themselves do not emit. This is highly suggestive that K3 was trained on Claude metadata (API logs, tagged synthetic data), rather than Claude's chat outputs. And it's well d…

For context you just asked me to read a document that starts with:

"Caveat: fully AI-generated research."

And that you quoted or paraphrased directly.

Re: The Kimi K3 Moment

#353

Earlier quoted context omitted.

> People infringe on Anthropics IP Unless someone literally stole the weights somehow (which is not out of the question, I doubt either oAI/Anthropic have the capabilities to prevent a state-level actor getting those weights), distillation from generations is not infringement on anyone's IP nor is it stealing nor is it an attack. It can't be. As long as you pay for tokens you get to do whatever you want with them. So…

Its definitely an attack. Thats established from anthropics perspective. No one has a right to use Anthropic’s services in ways that directly violate the ToS and user agreements.

> from anthropics perspective

I guess I can see that, if you mean the targeted effort of creating many accounts w/ the intent of doing it at scale. Sure, they may see that as an attack. But again, it's only an attack from their perspective if you agree that using generations to distill is "wrong". I just don't see it, in general. You can't both sell tokens and decide that distilling is somehow illegal. Something, something, cake and eat it.

Re: The Kimi K3 Moment

#354

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

The desire to accuse China of just copying is like 20 years out of date. It’s been wrong since some people on HN were in diapers. People are going to be gobsmacked when, in our lifetime, China becomes a world power comparable to the U.S. Probably still poorer per capita, but at Spain/Italy levels, not third world country levels. And they’ll be shocked at the implications of that on the world economy, migration patter…

There are huge evidence of copying.

Some day China can pioneer in science or technology but the current claim about Chinese companies leading AI development is ridiculous given the evidence of distillation and the fact that like 95 percent of science that lead to the current state of AI happened in either North America or Europe.

To be honest if you want to list academic papers that lead to the current AI models the majority is either done by Google Research or sponsored by Google.

Re: The Kimi K3 Moment

#355

Earlier quoted context omitted.

> 3.4 million is the number of sessions Anthropic detected. The actual number of Claude sessions trained on is likely >100 million. That's an increase of only a single order of magnitude, increasing my estimate of exfiltrated tokens from 0.05 to 0.15 trillion - a far cry from the 15 trillion required. > They are used for post-training Possibly - it may be too much data for post-training, unless further curation was d…

You're conflating pre-training data volume with post-training data volume. Nobody is suggesting Moonshot used 15 trillion tokens of Claude data to pre-train a base model from scratch. That would be impossible and nonsensical. This is entirely about distillation, which happens during post-training (alignment and SFT). Here, datasets are measured in millions or billions of tokens, not trillions. 50 billion Claude token…

There is a lot of supposition going on your part and mine. IMO, Chinese labs are not dependent on OpenAI/Anthropic outputs; they definitely use the outputs, but along other training/post-training data.

Now that Anthropic hides the real thinking tokens in a way that precludes future CoT distillation, we'll find out which side is correct based on whether Chinese AI labs close the gap or not.

My bet is they'll close the gap; nothing about frontier AI is magic, once something is shown to be possible, experienced practitioners almost always figure out how to accomplish the same feat, though not always on the same way. This is why frontier US labs keep leapfrogging each other every few months.

Re: The Kimi K3 Moment

#356

Earlier quoted context omitted.

China is infamous for weakly enforcing copyright law. Even when it is completely obvious that Chinese labs are training models on pirated data, US copyright holders face a virtually impossible task of proving it in court. Those lawsuits won't go anywhere.

There are tons of lawsuites which resulted in banning Chinese companies from doing business in US, those lawsuits totally have consequences.

What are the most high-profile examples of the "tons" of lawsuits resulting in Chinese companies being banned from doing business in the U.S.? Isn’t it usually more action by the government - executive orders, etc?

Re: The Kimi K3 Moment

#357

Earlier quoted context omitted.

the obvious difference is the massive scale of data and compute required to develop and evolve these models, and the costs they impose on those building them.

Smaller budgets, slower improvement, less risk. They're not entitled to profits if that business model isn't sustainable. They're not entitled to a change in IP laws to protect their business model. They're not entitled to growing that fast.

who are you talking about? again, my question is concerned with the "second class labs" and the sustainability of the distillation-as-a-service model.

Re: The Kimi K3 Moment

#358

Earlier quoted context omitted.

There's reproducible evidence of Kimi K3 spontaneously identifying itself as Claude https://x.com/denisewu/status/2077984660211269870 This behavior is exactly what you'd expect from a model distilled from Claude. Someone even took the time to analyze Kimi's ambiguous identity, in great detail: https://github.com/rgreenblatt/which_claude_is_k3/blob/main/... And there's an entire Reddit thread discussing this https://w…

By evidence I mean logs, I mean IP addresses, I mean timestamps. They claim millions of requests, let’s see literally any of them? I don’t consider a tweet by Denise Wu, who works at Anthropic, to be reproducible evidence. I don’t consider “Caveat: fully AI-generated research.” To be someone taking time to analyze anything in great detail. Because two AI models produce vaguely similar front-end styles when generating…

You've backtracked significantly here.

"Zero evidence" as you claimed earlier isn't accurate. You've moved the goalposts from "evidence" to "raw internal logs I can independently audit," which is a different and very high standard. Sure Anthropic didn't publish logs, IP addresses, timestamps, or account IDs of the accounts involved. But that's true of any cybersecurity breach/abuse disclosure ever made. Companies are furtive to reveal how they detect fraud, because doing so exposes the signals used to detect bad actors, and makes future abuse easier. Not revealing the "evidence" you're asking for is industry standard practice. You're complaining that Anthropic is following industry standard practice, and conveniently defining the "evidence" you need as something Anthropic is never going to publish.

> I don’t consider a tweet by Denise Wu, who works at Anthropic, to be reproducible evidence.

Is the issue here that she works at Anthropic? Because Denise Wu doesn't work there.

> I don’t consider “Caveat: fully AI-generated research” to be someone taking time to analyze anything in great detail.

The experiments were run by Ryan Greenblatt, who is a real AI safety researcher (at Redwood Research).

The identity experiments and Greenblatt analysis are trivially reproducible. The methodology, code, and metrics are all there in the Github repository. You can ask your preferred AI to independently replicate these results, and it will give you a result within an hour.

You’ve also reduced the evidence to “two models producing vaguely similar front-end styles,” which is not what either analysis shows.

From the analysis, Kimi K3 identifies itself as Claude 15% of the time. How do you explain that? Qwen and GPT identify themselves as Claude 0% of the time.

If a long document is too much analysis for you, someone else made a simple chart which measures the KL divergence between Kimi K3 and other major models. They found K3 is unusually similar to Fable 5 & Opus models. That is, Kimi K3 has an very similar style and phrasing to that of Anthropic models. That behavior is expected from a model distilled from Claude.

https://typebulb.com/u/lab/you-re-relatively-right/full

Re: The Kimi K3 Moment

#359

Earlier quoted context omitted.

I strongly agree with the premise that distillation is not an “attack”. But that said: K3 is not a distilled version of Fable or Sol. Fable has been barely available and Sol was just released! Moreover, K3 is superior to both models in some domains, according to user scoring on the Arena. API distillation can’t give you these results anyway. All it is useful for is bootstrapping RL in new domains to get past the “col…

API distillation doesn't have to explain all of K3's capabilities for it to have happened. Kimi K3 reproducibly identifies itself as Claude: https://x.com/denisewu/status/2077984660211269870 This behavior is exactly what you'd expect from a model distilled from Claude. There's a detailed analysis of K3's ambiguous identity here: https://github.com/rgreenblatt/which_claude_is_k3/blob/main/... This analysis observed K3…

It does not reproducibly identify itself as Claude, there's evidence to the contrary in the very thread you linked: https://x.com/bobbyNewcomb5/status/2078151562828947954

Re: The Kimi K3 Moment

#360

Earlier quoted context omitted.

API distillation doesn't have to explain all of K3's capabilities for it to have happened. Kimi K3 reproducibly identifies itself as Claude: https://x.com/denisewu/status/2077984660211269870 This behavior is exactly what you'd expect from a model distilled from Claude. There's a detailed analysis of K3's ambiguous identity here: https://github.com/rgreenblatt/which_claude_is_k3/blob/main/... This analysis observed K3…

It does not reproducibly identify itself as Claude, there's evidence to the contrary in the very thread you linked: https://x.com/bobbyNewcomb5/status/2078151562828947954

As mentioned in my comment, Kimi K3 identifies itself as Claude ~15% of the time.

Here's another report of K3 identifying itself as Claude https://x.com/Sauers_/status/2077842686459981901

And an analysis showing the self-identity distribution for K3 and other models https://x.com/RyanGreenblatt/status/2078663148509544589

Post reply on HN