Earlier quoted context omitted.
Sure, there is litigation, criminal case, appeals, fines ($500M: https://www.motorolasolutions.com/newsroom/press-releases/hy... ). The point is if violation is clear, US corps have a chance to go after Chinese corps.
>> There are tons of lawsuites which resulted in banning Chinese companies from doing business in US > What are the most high-profile examples of the "tons" of lawsuits resulting in Chinese companies being banned from doing business in the U.S.? Isn’t it usually more action by the government - executive orders, etc? In response to "What are the most high-profile examples of lawsuits resulting in Chinese companies bei…
The Kimi K3 Moment
471–480 of 644 posts
Re: The Kimi K3 Moment
#472Earlier quoted context omitted.
While it sounds like a lot, do you suppose 3.4 million sessions come even close to being sufficient to train a frontier model? Assuming each session was 10,000 words each, that's 34 billion words; lets call it 50 billion tokens (0.05 trillion) unfairly pilfered from Claude. That left Moonshot needing to scrounge for the other 14.950 trillion training tokens required for a baseline frontier model.
What do you think those tokens are used for? Distillation attacks aren't about replacing the entire pretraining dataset with questionably sourced synthetics. It's all about post-training. Train your own base model - but tune it off Claude output to make it perform more in line with Claude. Yoink the products of Anthropic's expensive SFT, RLHF and RLVR work for yourself by training on the outcomes. The post-training d…
Is that actually genuine distillation though? Distillation suggests the core model is being pre-trained using output from another model. For the above to work, you have to already have all the core intelligence trained into your base model.
If distillation just comes down to post-training then it's tantamount to admitting that the Chinese base models are just as good as frontier US lab models. Because you can't post-train frontier intelligence into a model. It has to be there in the base. Then you can change how that intelligence is expressed through post-training.
Re: The Kimi K3 Moment
#473Earlier quoted context omitted.
TOS violations are not espionage. Everybody who links up Claude to OpenCode is violating the TOS. So from the largest industrial espionage in history we have left "They paid for the accounts but violated the TOS". And then you randomly add the claim they stole the money to pay for the accounts. You have provided no evidence other than "Claude tokens are sold for cheap in China". As others have pointed out, that might…
Have a read of this detailed article, it's well sourced and documented that token resellers are logging the Claude outputs and selling them to Chinese labs. All your points are addressed in there https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens... I linked it earlier, but it seems you didn't see it. re: labs purchasing model/tool output, see https://x.com/xkajon/status/2050445443889525235 re: model swappi…
Second, I now actually read the article. It describes plenty of questionable and problematic things but also contradicts your claims explicitly.
The essential point of the article is about making money by selling access to Claude cheaply in China. Not about Chinese labs orchestrating a way to get their hands on Claude output.
Your credit card claim is considerably weaker in the article: "[Beyond this there are] accounts purchased using stolen or fraudulent credit cards [...]. How large this share is relative to the above four “innocent” tactics is difficult to verify, but the two markets likely share some infrastructure and personnel.". Instead, swapping models to cheaper alternatives is listed as a major reason for cheaper prices.
Then the article gets a key point wrong: As many others have pointed out to you, you don't get access to the reasoning traces anymore on the subscription accounts. And the article also clearly states:
"Chinese developer communities assert [selling logs] is happening in at least some cases, but whether proxy operators are systematically harvesting and selling these logs, and to whom, remains unverified. However, downstream distillation data does exist on the open web. Several datasets of Claude Opus 4.6 reasoning outputs circulate on HuggingFace with no clear source for the outputs. Theoretically, one can clean and sell similar distilled datasets to other model developers in China."
The article also discusses selling logs for other (far worse!) purposes than for training, like blackmail.
So overall this article reads very, very different to your claims. Nothing in the article suggests or supports the idea of large-scale coordinated "distillation attacks". Instead it paints the picture of a naturally emerging grey-market response to access control blocks, consisting of many exchangeable individual actors: "Almost no one operates the full chain. Most participants own one or two links and monetise those well, resulting in a resilient, modular system."
Importantly: Nothing in any of this looks ethically worse to me than Meta using pirated books for training. And nothing suggests that OpenAI or Anthropic were more ethical than Meta when sourcing their material.
Re: The Kimi K3 Moment
#474Earlier quoted context omitted.
TOS violations are not espionage. Everybody who links up Claude to OpenCode is violating the TOS. So from the largest industrial espionage in history we have left "They paid for the accounts but violated the TOS". And then you randomly add the claim they stole the money to pay for the accounts. You have provided no evidence other than "Claude tokens are sold for cheap in China". As others have pointed out, that might…
Have a read of this detailed article, it's well sourced and documented that token resellers are logging the Claude outputs and selling them to Chinese labs. All your points are addressed in there https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens... I linked it earlier, but it seems you didn't see it. re: labs purchasing model/tool output, see https://x.com/xkajon/status/2050445443889525235 re: model swappi…
Re: The Kimi K3 Moment
#475Earlier quoted context omitted.
The main story is what isn't being talked about. Chinese labs exfiltrated trillions of tokens of high-quality output from Anthropic and OpenAI, through proxies and heavily discounted token resellers, which they distilled and used for training data for their own models. Instead of spending 12-18 months building their own robust harnesses and painstakingly creating quality training data (which is what Anthropic and Ope…
I could perhaps get myself to care just the tiniest bit if the information that was supposedly stolen wasn't generated by "stealing" from everybody else. Either it is fair use to train AI models on whatever information you can get your hands on for everyone or for no one.
Re: The Kimi K3 Moment
#476Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…
Re: The Kimi K3 Moment
#477Even in this very thread the feedback on Kimi's actual efficacy is debated. I personally feel its worse than both Fable and 5.6 Sol, but I feel like the conversation isn't really about whether its good or not, but a backlash against the U.S governments foray into regulation. So I think people _want_ it to be superior out of anger/frustration with the current situation.
This seems like a replay of what happened with DeepSeek. They put out v3, or whichever one it was, and everyone said it was over for US companies... then everything continued on.
And now 100% to a mix of K3 / DeepSeek V4 / MiMo 2.5.
It's nice not being called a terrorist just because I told it to reverse engineer something.
At work they are still hemorrhaging money to Western providers due to enterprise contracts but I foresee they won't renew for much longer. Specially of the upcoming final version of DeepSeek V4 proves to be Opus+ level.
Re: The Kimi K3 Moment
#478Earlier quoted context omitted.
What do you think those tokens are used for? Distillation attacks aren't about replacing the entire pretraining dataset with questionably sourced synthetics. It's all about post-training. Train your own base model - but tune it off Claude output to make it perform more in line with Claude. Yoink the products of Anthropic's expensive SFT, RLHF and RLVR work for yourself by training on the outcomes. The post-training d…
> Train your own base model - but tune it off Claude output to make it perform more in line with Claude Is that actually genuine distillation though? Distillation suggests the core model is being pre-trained using output from another model. For the above to work, you have to already have all the core intelligence trained into your base model. If distillation just comes down to post-training then it's tantamount to ad…
You have to bring those bits and pieces together, put them into the right shapes and fill in the gaps to get a model that actually performs. This is what post-training is all about. It's not at all a trivial thing.
Reasoning, tool use, agentic behavior - all of those are post-training performance gains. Getting a good well trained base model is putting your foot in the door of frontier performance - post-training is how you actually get inside.
See: GPT-4.5 vs o1. One went for "build a bigger better more capable base model", the other went for "take the old base and post-train it for advanced capabilities". The results: a wider base with basic post-training loses to a narrower base with advanced post-training. Or, hell: GPT-3 vs GPT-3.5. One was largely a research lab curio, and the other kicked off the AI revolution as we know it.
The gains compound. Getting a better base model with the same type of post-training helps, see: the jump from Opus to Mythos/Fable. But post-training techniques account for a lot of the performance juice.
And yes, reasoning trace post-training distillation is "genuine distillation". As is logit distillation in pre-training. "Distillation" isn't a single training recipe that you have to follow to a tee - it's a large group of training methods. I've seen plenty of wacky things like inverse distillation bootstrap and post-training self-distillation that use distillation in strange ways at different stages of the training run to get results.
Re: The Kimi K3 Moment
#479Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…
The desire to accuse China of just copying is like 20 years out of date. It’s been wrong since some people on HN were in diapers. People are going to be gobsmacked when, in our lifetime, China becomes a world power comparable to the U.S. Probably still poorer per capita, but at Spain/Italy levels, not third world country levels. And they’ll be shocked at the implications of that on the world economy, migration patter…
Re: The Kimi K3 Moment
#480Earlier quoted context omitted.
if Ford bought hundreds of millions of dollars worth of Hyundais, put extra instrumentation in them, and resold them at a discount to customers who agreed to the instrumentation in exchange for the discount, is Ford doing industrial espionage?
You skipped the part where Ford buys the cars at 90% off and sells them at 80% off, at a profit. Then gets paid by competitors for the driving data. At the same time, Volvo is running the exact same hustle, except they buy the cars with stolen credit cards, so they get the cars for free.