Live data from Hacker News

The Kimi K3 Moment

stephen.bochinski.dev

401–410 of 644 posts

Re: The Kimi K3 Moment

#402

Earlier quoted context omitted.

API distillation doesn't have to explain all of K3's capabilities for it to have happened. Kimi K3 reproducibly identifies itself as Claude: https://x.com/denisewu/status/2077984660211269870 This behavior is exactly what you'd expect from a model distilled from Claude. There's a detailed analysis of K3's ambiguous identity here: https://github.com/rgreenblatt/which_claude_is_k3/blob/main/... This analysis observed K3…

> Kimi K3 reproducibly identifies itself as Claude It could also be have been trained from collected response datasets. Claude got caught several time responding it was ChatGPT or even Deepseek and I don't think Anthropic has been distealling DeepSeek. > This behavior is exactly what you'd expect from a model distilled from Claude. The opposite actually. If they wanted to distill Claude without getting caught they co…

  > distealling
Apt typo.

Though I am of the opinion that distilling is no different than how extant frontier LLMs have also been trained on other people's data, I could actually see the word distealling becoming useful in discussion.

Re: The Kimi K3 Moment

#403

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

The hand wringing over whether internationally located AI labs are "stealing" output from American ones is the funniest thing in a while.

It's international politics with people talking about AI success as a matter of national strategic advantage and survival. So at best "this was built off our work" mostly tells you that apparently you've got months of advantage when a new model drops before it can be cloned. That's certainly some sort of advantage, sure hope it represents a consistent ability to stay ahead and causes people to redouble their efforts.

Or...of course none of these companies are worth what they say, but the advantage is also not really that great, and a whole lot of people are just really worried about their stock payouts.

Re: The Kimi K3 Moment

#404

Earlier quoted context omitted.

I strongly agree with the premise that distillation is not an “attack”. But that said: K3 is not a distilled version of Fable or Sol. Fable has been barely available and Sol was just released! Moreover, K3 is superior to both models in some domains, according to user scoring on the Arena. API distillation can’t give you these results anyway. All it is useful for is bootstrapping RL in new domains to get past the “col…

API distillation doesn't have to explain all of K3's capabilities for it to have happened. Kimi K3 reproducibly identifies itself as Claude: https://x.com/denisewu/status/2077984660211269870 This behavior is exactly what you'd expect from a model distilled from Claude. There's a detailed analysis of K3's ambiguous identity here: https://github.com/rgreenblatt/which_claude_is_k3/blob/main/... This analysis observed K3…

Kimi calling itself claude means nothing. During pre-training, when the model learns to "simulate" the internet text, it will naturally be fed with a bunch of data about Claude and ChatGPT. With the amount of LLM outputs on the internet today, it is not surprising at all that a model would naturally call itself Claude or ChatGPT. You can mitigate that in post-training (or actually in pre-training as well) by training on many examples of what the model should call itself. That being said, getting probably hundreds pf thousands of ChatGPT and Claude examples totally "pirged" out of the weights is going to be difficult and really more hassle than its worth.

Re: The Kimi K3 Moment

#405
post #335

Earlier quoted context omitted.

> western governments Are you talking about the US, specifically? Why would other countries, that don't share the same anxiety about China as the US, would be troubled with the this?

"Why would other countries, that don't share the same anxiety about China as the US, would be troubled with the this?" It's the other way around. There is a high likelihood that many countries of the "west" (the "global north"?) will outlaw, restrict, or otherwise control LLMs and the tools that enable them. The US, however, is blessed with the first amendment which makes it extremely difficult to restrain speech in…

[flagged]

Re: The Kimi K3 Moment

#406
post #335

Earlier quoted context omitted.

> western governments Are you talking about the US, specifically? Why would other countries, that don't share the same anxiety about China as the US, would be troubled with the this?

"Why would other countries, that don't share the same anxiety about China as the US, would be troubled with the this?" It's the other way around. There is a high likelihood that many countries of the "west" (the "global north"?) will outlaw, restrict, or otherwise control LLMs and the tools that enable them. The US, however, is blessed with the first amendment which makes it extremely difficult to restrain speech in…

It wasn't difficult for the US to restrict TikTok (or BYD, Huawei, DJI...)

Re: The Kimi K3 Moment

#407

Earlier quoted context omitted.

The desire to accuse China of just copying is like 20 years out of date. It’s been wrong since some people on HN were in diapers. People are going to be gobsmacked when, in our lifetime, China becomes a world power comparable to the U.S. Probably still poorer per capita, but at Spain/Italy levels, not third world country levels. And they’ll be shocked at the implications of that on the world economy, migration patter…

There are huge evidence of copying. Some day China can pioneer in science or technology but the current claim about Chinese companies leading AI development is ridiculous given the evidence of distillation and the fact that like 95 percent of science that lead to the current state of AI happened in either North America or Europe. To be honest if you want to list academic papers that lead to the current AI models the…

Okay but I cannot stress this enough: no one cares.

It's international politics. The rules are optional, and written on the back of whoever agrees to enforce them.

If you're going to run around declaring AI is a strategic advantage vital to national security, then guess what? Stealing it is a great idea. That you stole it is only a problem if it means you're not developing the ability to support that work locally as well, and China seems to be doing very well at building it's local talent and support network.

If you ever listen to Russian propaganda, there's a similar theme: every big idea, everything good, all of it was definitely first developed in Russia - only Russians could ever have thought of it. Of course, Russia isn't actually a world leader in any of those things, or able to execute on them.

Which is what America is sounding like more and more these days.

Re: The Kimi K3 Moment

#408
post #264

Earlier quoted context omitted.

Yes, anthropic and openai have really been brought to their knees and ipos cancelled because of the legal consequences of obtaining their training data.

This would have been a problem but it turns out that Anthropic is actually valued multiple orders of magnitude more than a copy of all the books in the world. So they survived the significant legal consequences.

In a just world, the punishment can be more than just the sum of the direct damages, otherwise there's no incentive to stop reoffending.

Anthropic (and friends) proved they're willing to do obviously illegal things, and it didn't end them. Why do we think they stopped after doing it once?

Re: The Kimi K3 Moment

#409
It’s still >$300k for the hardware to run this model locally at anything resembling a reasonable speed. The weights also have not yet been released, though that is scheduled for about a week from now. It’s not actually an open model yet.

Re: The Kimi K3 Moment

#410

Earlier quoted context omitted.

Shrug. Hard to feel sorry for companies that created their empires by ignoring copyright themselves. Also, 'most extensive industrial espionage campaign, probably ever' is absolute nonsense. They did not need to infiltrate the companies for this nor are you accusing them of stealing any trade secrets. This is only about whether they looked at their competitors' products from the outside (in the form of conversation t…

> 'most extensive industrial espionage campaign, probably ever' is absolute nonsense You are completely underestimating the scale of what is happening here. Chinese AI labs are actively facilitating an industrial-scale network of tens of thousands of bot accounts, that resell Claude tokens at 97% below official API prices. They buy subsidized Max 5x plans (sometimes with stolen credit cards), then split the subscript…

You haven't explained how this is illegal or any more immoral than scraping the web for training data.

As you said yourself: They are buying the product. Then they are using it for their own purposes. That's more than Anthropic/OpenAI did for the open internet. That's more than Meta did when they obtained torrents of books in the early days, and then claimed that even though the data was obtained illegally they can still train on it just fine.

They paid for it! It's absurd to call this espionage!

Post reply on HN