Live data from Hacker News

The Kimi K3 Moment

stephen.bochinski.dev

441–450 of 644 posts

Re: The Kimi K3 Moment

#441

Earlier quoted context omitted.

By evidence I mean logs, I mean IP addresses, I mean timestamps. They claim millions of requests, let’s see literally any of them? I don’t consider a tweet by Denise Wu, who works at Anthropic, to be reproducible evidence. I don’t consider “Caveat: fully AI-generated research.” To be someone taking time to analyze anything in great detail. Because two AI models produce vaguely similar front-end styles when generating…

You've backtracked significantly here. "Zero evidence" as you claimed earlier isn't accurate. You've moved the goalposts from "evidence" to "raw internal logs I can independently audit," which is a different and very high standard. Sure Anthropic didn't publish logs, IP addresses, timestamps, or account IDs of the accounts involved. But that's true of any cybersecurity breach/abuse disclosure ever made. Companies are…

"From the analysis, Kimi K3 identifies itself as Claude 15% of the time. How do you explain that? Qwen and GPT identify themselves as Claude 0% of the time."

Qwen and GPT have special guards that trigger when asked to identify, Kimi doesnt. I dont understand the argument. Kimi is an LLM and does not know what it is. It will give you the most likely answer which sometimes is Claude.

Re: The Kimi K3 Moment

#442
post #407

Earlier quoted context omitted.

There are huge evidence of copying. Some day China can pioneer in science or technology but the current claim about Chinese companies leading AI development is ridiculous given the evidence of distillation and the fact that like 95 percent of science that lead to the current state of AI happened in either North America or Europe. To be honest if you want to list academic papers that lead to the current AI models the…

Okay but I cannot stress this enough: no one cares. It's international politics. The rules are optional, and written on the back of whoever agrees to enforce them. If you're going to run around declaring AI is a strategic advantage vital to national security, then guess what? Stealing it is a great idea . That you stole it is only a problem if it means you're not developing the ability to support that work locally as…

> If you ever listen to Russian propaganda, there's a similar theme: every big idea, everything good, all of it was definitely first developed in Russia - only Russians could ever have thought of it. Of course, Russia isn't actually a world leader in any of those things, or able to execute on them.

When I was a kid watching Star Trek VI, I was confused by the line "You've not experienced Shakespeare until you've read him in the original Klingon".

And then I learned about how the Klingons (especially in that film) were a stand-in for the USSR.

Re: The Kimi K3 Moment

#443

Earlier quoted context omitted.

You haven't explained how this is illegal or any more immoral than scraping the web for training data. As you said yourself: They are buying the product. Then they are using it for their own purposes. That's more than Anthropic/OpenAI did for the open internet. That's more than Meta did when they obtained torrents of books in the early days, and then claimed that even though the data was obtained illegally they can s…

> They paid for it! They didn't though. The resellers are not buying via the official API, they're buying Max subscriptions (where tokens are priced ~10x below API cost), then splitting the subscription across dozens of clients and reselling the output as the regular API. Anthropic prices its subscription plans barely at cost, to bring in customers onto their enterprise plans where they can charge expensive API rates…

TOS violations are not espionage. Everybody who links up Claude to OpenCode is violating the TOS.

So from the largest industrial espionage in history we have left "They paid for the accounts but violated the TOS". And then you randomly add the claim they stole the money to pay for the accounts.

You have provided no evidence other than "Claude tokens are sold for cheap in China". As others have pointed out, that might also simply be counterfeit tokens generated by open weight models.

The western labs have established the precedent that all data they can buy beg borrow or steal is fair game. Turning around and crying foul when the Chinese labs follow their lead is hypocrisy.

Re: The Kimi K3 Moment

#444
post #442
post #407

Earlier quoted context omitted.

Okay but I cannot stress this enough: no one cares. It's international politics. The rules are optional, and written on the back of whoever agrees to enforce them. If you're going to run around declaring AI is a strategic advantage vital to national security, then guess what? Stealing it is a great idea . That you stole it is only a problem if it means you're not developing the ability to support that work locally as…

> If you ever listen to Russian propaganda, there's a similar theme: every big idea, everything good, all of it was definitely first developed in Russia - only Russians could ever have thought of it. Of course, Russia isn't actually a world leader in any of those things, or able to execute on them. When I was a kid watching Star Trek VI, I was confused by the line "You've not experienced Shakespeare until you've read…

But that was more of a jab at literary snobs who would tout "Homer in the original Greek" or "Marcus Aurelius in the original Latin" or "Old Testament in the Original Hebrew". It has been such a meme, probably for centuries. Because it was not so long ago when university students were actually conversant in many classical languages such as those.

Re: The Kimi K3 Moment

#445

Earlier quoted context omitted.

API distillation doesn't have to explain all of K3's capabilities for it to have happened. Kimi K3 reproducibly identifies itself as Claude: https://x.com/denisewu/status/2077984660211269870 This behavior is exactly what you'd expect from a model distilled from Claude. There's a detailed analysis of K3's ambiguous identity here: https://github.com/rgreenblatt/which_claude_is_k3/blob/main/... This analysis observed K3…

and claude will call itself chatgpt etc. nothing new, all ai labs are immoral and not bound by any reasonable oversight or ethical constraints. All outlaws in their own rights on that front. Absolutely none of them have true rights on the matter of being distilled from given historic and continued behaviour. I'm not sure why this is a talking point at all? We know AI companies steal, the least interesting behaviour a…

> For me, a far more interesting and important point of conversation on this matter is anthropic buying rare or evwn unique books, processing them for training data, and then destroying the books for others cannot use it as well.

That's an incredible allegation, and appalling if true. But is it true?

Re: The Kimi K3 Moment

#446
post #335

Earlier quoted context omitted.

> western governments Are you talking about the US, specifically? Why would other countries, that don't share the same anxiety about China as the US, would be troubled with the this?

"Why would other countries, that don't share the same anxiety about China as the US, would be troubled with the this?" It's the other way around. There is a high likelihood that many countries of the "west" (the "global north"?) will outlaw, restrict, or otherwise control LLMs and the tools that enable them. The US, however, is blessed with the first amendment which makes it extremely difficult to restrain speech in…

> There is a high likelihood that many countries of the "west" (the "global north"?) will outlaw, restrict, or otherwise control LLMs and the tools that enable them.

Why? Based on what?

I've seen absolutely no indication of this.

Only the US are playing this game atm.

Re: The Kimi K3 Moment

#447

Earlier quoted context omitted.

Kimi calling itself claude means nothing. During pre-training, when the model learns to "simulate" the internet text, it will naturally be fed with a bunch of data about Claude and ChatGPT. With the amount of LLM outputs on the internet today, it is not surprising at all that a model would naturally call itself Claude or ChatGPT. You can mitigate that in post-training (or actually in pre-training as well) by training…

Sure, but then Qwen should leak that too, and it doesn't. K3 calls itself Claude 7 out of 48 times, Qwen does it 0 out of 48, and the only other model to identify itself as Claude is DeepSeek. and DeepSeek is alleged to also distill from Claude data anyway. So this isn't something every model absorbed from the same web text. And you skipped over the strongest datapoint that K3 is distilled: K3 reproduces Claude's pub…

https://qwen.readthedocs.io/en/latest/training/ms_swift.html

Qwen cares enough about model identity that their training framework and docs include a preset for training on it complete with a targeted dataset: https://huggingface.co/datasets/modelscope/self-cognition

And people get Claude to claim it's Deepseek by asking in Chinese.

I can't believe we're still at the "I asked the model who it is" stage of LLMs nearly 4 years out from models calling themselves GPT by OpenAI.

Re: The Kimi K3 Moment

#448

Earlier quoted context omitted.

What do you think those tokens are used for? Distillation attacks aren't about replacing the entire pretraining dataset with questionably sourced synthetics. It's all about post-training. Train your own base model - but tune it off Claude output to make it perform more in line with Claude. Yoink the products of Anthropic's expensive SFT, RLHF and RLVR work for yourself by training on the outcomes. The post-training d…

How does yoinking outputs from from prior generation Claude model and post raining on them result in a model competitive with the latest generation? That doesn't add up - nevermind Anthropic hasbeen summarizing thinking tokens since January to counter distillation.

Do I really have to explain the shape of AI training pipelines to you?

Train a big, wide base model with a lot of potential. Mid-train or post-train that on Claude Opus 4.5 reasoning/agentic traces (i.e. Claude Code data from Chinese API resellers) to make your model approximate a high baseline of chatbot behavior, reasoning, agentic work and tool use.

Then run your own expensive SFT, RLHF and RLVR on top of that yoinked baseline to dial it in further.

Actually doing RLHF and RLVR is extremely expensive. Distillation gives you a lot of dense, high quality post-training signal for cheap. This can get your model into the basin of "the right way to tackle this kind of problem" without a frontier lab compute budget. It's a big shortcut that gets you closer to the target - you can take it from there and build on top of it with your own work.

Also, it's unclear whether "summarizing thinking tokens" actually ruins distillation, or just makes it work worse. I'd bet on the latter, really. Because it's an approximation game, and summarized reasoning is still a better approximation of true reasoning than most of what you get online and in pre-training datasets.

Re: The Kimi K3 Moment

#449

> I’ve been running Kimi K3 alongside Claude on my normal coding work, and for all practical purposes I can’t tell them apart When you say "Claude", do you mean Opus? Fable? What effort level?

This is comparing Fable High with K3 High. I'm mostly using these models for game development. The tasks I usually send are ambiguous visual bugs, changing the look of a scene or models, or adding a large feature. The wording wasn't accurate there. I don't use Fable or K3 most of the time. I'm usually working on smaller scoped tasks that I review myself afterwards.

Please update the blog post to clarify. “Claude” is not a model, and your writing makes no sense without specifying.

Re: The Kimi K3 Moment

#450

Earlier quoted context omitted.

and claude will call itself chatgpt etc. nothing new, all ai labs are immoral and not bound by any reasonable oversight or ethical constraints. All outlaws in their own rights on that front. Absolutely none of them have true rights on the matter of being distilled from given historic and continued behaviour. I'm not sure why this is a talking point at all? We know AI companies steal, the least interesting behaviour a…

> For me, a far more interesting and important point of conversation on this matter is anthropic buying rare or evwn unique books, processing them for training data, and then destroying the books for others cannot use it as well. That's an incredible allegation, and appalling if true. But is it true?

It's not an allegation https://www.washingtonpost.com/technology/2026/01/27/anthrop... (if you're talking about the "rare or unique" part, yeah that might be bs)

But in my opinion, treating mass produced books like they're this sacred untouchable object is ridiculous. They're not "source" material, they're just a copy as well, and they're not "priceless" by any means. They're very reasonably priced, perhaps even so cheaply priced that books can be bought in bulk in these amounts. Buying used books and doing whatever you want with them is just legal. Used books, that would probably be just laying in some warehouse, or recycled anyway.

If there's anything to have gripes with, it's the copyright system that makes it easier to take this legal route.

Post reply on HN