Live data from Hacker News

The Kimi K3 Moment

stephen.bochinski.dev

601–610 of 644 posts

Re: The Kimi K3 Moment

#601

Earlier quoted context omitted.

> 3.4 million is the number of sessions Anthropic detected. The actual number of Claude sessions trained on is likely >100 million. That's an increase of only a single order of magnitude, increasing my estimate of exfiltrated tokens from 0.05 to 0.15 trillion - a far cry from the 15 trillion required. > They are used for post-training Possibly - it may be too much data for post-training, unless further curation was d…

You're conflating pre-training data volume with post-training data volume. Nobody is suggesting Moonshot used 15 trillion tokens of Claude data to pre-train a base model from scratch. That would be impossible and nonsensical. This is entirely about distillation, which happens during post-training (alignment and SFT). Here, datasets are measured in millions or billions of tokens, not trillions. 50 billion Claude token…

> I don't understand how you're so caught up on the term "distillation". Distillation is using a larger model's outputs to train a (weaker) student model. Which is exactly what's happening.

But from evidence it was data generated by Opus 4.5. Is Opus 4.5 larger, stronger model and Kimi K3 a weaker, student model? I don't think so.

Also, I searched on HuggingFace "fable data", first link was a "distillation dataset", including samples from Fable, there are 2 milion code samples of Fable output. So this fully legal and sits in the open, somehow is not a distillation attack? Who knows, maybe Moonshot used this for training.

Re: The Kimi K3 Moment

#602

Earlier quoted context omitted.

my first prompt to any Kimi model was K3 via Pi, some version of "hi kimi!!" and the response was telling me "I'm actually Claude." this is not hard to repro, just use a system prompt that doesn't mention the model name. that said, if they bootstrapped with opus 4.6 convo sft data they had sitting around... so what?

The main story is what isn't being talked about. Chinese labs exfiltrated trillions of tokens of high-quality output from Anthropic and OpenAI, through proxies and heavily discounted token resellers, which they distilled and used for training data for their own models. Instead of spending 12-18 months building their own robust harnesses and painstakingly creating quality training data (which is what Anthropic and Ope…

>The community is happy to overlook any questionable methods by Chinese Labs.

Using the US-based models are arguably even more questionable. You have to be content with the OpenAI and Anthropic literally scraping the entire internet. They've all pirated content, scraped against ToS, ignored robots.txt, bypassed paywalls all to train their models. It's well known these AI labs have ingested the entirety of Annas-Archive into their models, the largest collection of books ever assembled.

They didn't credit or compensate literally any artist, author, scientist or publicist in the creation of their models.

They've lobbied against local Governments to shove big and loud datacentres in peoples backyard. They've polluted local water supplies, they've doubled energy costs for these regions. They've tormented locals with subsonic frequencies.

They've given access to the DoD to use their models to kill people, or assist in killing people. They've used these models to enable mass surveillance, allegedly not domesicially but since when can we trust any of the 3-letter agencies.

Using "US Models" is not the moral high ground you think it is. Kimi saying its Claude, Gemini or ChatGPT is not the "substantial evidence" you think either.

>We witnessed the most extensive industrial espionage campaign, probably ever, and nobody in the industry cares at all that it happened.

Because they stole for every one of us without permission. Thousands of my comments on this site and others (Stackoverflow, etc) are all used in their training data.

Not to mention, OpenAI has allegedly just stole tons of internal Apple documents... I guess we'll just ignore that too.

Re: The Kimi K3 Moment

#603

Earlier quoted context omitted.

Have a read of this detailed article, it's well sourced and documented that token resellers are logging the Claude outputs and selling them to Chinese labs. All your points are addressed in there https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens... I linked it earlier, but it seems you didn't see it. re: labs purchasing model/tool output, see https://x.com/xkajon/status/2050445443889525235 re: model swappi…

First off, I don't doubt that China engages in large-scale industrial espionage, or at least used to. Nowadays they have plenty of talented and highly educated engineering talent of their own. Second, I now actually read the article. It describes plenty of questionable and problematic things but also contradicts your claims explicitly. The essential point of the article is about making money by selling access to Clau…

> Nothing in the article suggests or supports the idea of large-scale coordinated "distillation attacks"

Did you read the correct article? This is covered in the first paragraph:

https://www.anthropic.com/news/detecting-and-preventing-dist...

It directly addresses large-scale, coordinated 'distillation attacks' orchestrated by Chinese labs, which Anthropic accuses of exfiltrating tens of millions of exchanges. The rest of the article elaborates on how this is done. Specifically, how these labs use transfer stations mix in genuine user traffic to conceal the distillation.

> Your credit card claim is considerably weaker in the article

I only mentioned payments fraud because you asserted without evidence that Anthropic is actually getting paid for these tokens.

You wrote, without providing any sources, that these proxy services are "buying the product" and that "They paid for it!"; and used that as justification for their behavior.

My counterpoint directly invalidates that assumption. While perhaps not every single reseller relies on fraud, you blindly generalized that these proxy services are all legitimate, paying customers.

In reality, the industry exists on shady practices. As detailed by this industry insider https://x.com/yan5xu/status/2029743983522631698 , these operations routinely:

* Use botnets to mass-create thousands of accounts

* Blatantly violate ToS by splitting and reselling account access

* Use fraudulent identities to create thousands of bot accounts

* Bypass KYC by recruiting real people in low-income countries for biometric face-matching checks for a few dollars

* Use AI deepfakes to fake passports / verification credentials

You can't claim they "They paid for it!" when the entire system is built on systematic fraud.

Re: The Kimi K3 Moment

#604

Even in this very thread the feedback on Kimi's actual efficacy is debated. I personally feel its worse than both Fable and 5.6 Sol, but I feel like the conversation isn't really about whether its good or not, but a backlash against the U.S governments foray into regulation. So I think people _want_ it to be superior out of anger/frustration with the current situation.

I can't count the number of times I've heard people here say the frontier models are 6 months or more ahead of the open-weights models. That's not true anymore. So the goalposts are shifting.

I feel like that was true up until 6 months ago.

Re: The Kimi K3 Moment

#605

Earlier quoted context omitted.

This is comparing Fable High with K3 High. I'm mostly using these models for game development. The tasks I usually send are ambiguous visual bugs, changing the look of a scene or models, or adding a large feature. The wording wasn't accurate there. I don't use Fable or K3 most of the time. I'm usually working on smaller scoped tasks that I review myself afterwards.

Please update the blog post to clarify. “Claude” is not a model, and your writing makes no sense without specifying.

There’s probably some other things that weren’t exactly clear or could be misinterpreted in the post. I don’t see it as worth my time to go back and edit each little thing.

Re: The Kimi K3 Moment

#606

Earlier quoted context omitted.

I mean, you have those kind of luxuries even in the poorest of countries in Asia, it’s just that there’s still a huge discrepancy between rich and poor, city vs countryside. It’s not difficult to find areas in all these countries that are significantly less developed than Spain/Portugal’s underdeveloped areas. It’s just not as black and white as you seem to suggest. (I come from EU but have been living in various cou…

From what I can tell, China has a capitalism problem, not an industrial wealth problem - they could house every one of their citizens in a decent apartment and give them gadgets and a car to drive, yet half the country is third-world poor still. But that's because in a market economy, you can't just give away things, and those poor people don't produce much of value by capitalist standards, so the can't pay for those…

> From what I can tell, China has a capitalism problem, not an industrial wealth problem - they could house every one of their citizens in a decent apartment and give them gadgets and a car to drive

This is wildly incorrect.

Re: The Kimi K3 Moment

#607

Earlier quoted context omitted.

China never allow US AI in China, so they HAVE to build Chinese equivalents... US immigration policy isn't a big factor. China's got 1.8B people. If you don't think they've got the talent to pull this off, even if a lot of it leaves to live elsewhere, you're naive. No one uses Baidu, but they built their own Google, and it's good. They built their own Facebooks and Instagrams. The US isn't the only place in the world…

"China can draw on a talent pool of 1.3 billion people, but the United States can draw on a talent pool of 7 billion and recombine them in a diverse culture that enhances creativity in a way that ethnic Han nationalism cannot." --Lee Kuan Yew, former prime minister of Singapore.

This quote feels pretty delusional even at the time (2015), although 11 years later, this quote (view) holds almost no water imo.

Re: The Kimi K3 Moment

#608
post #466

Earlier quoted context omitted.

In this world too, the punishment was, ballpark, 100x direct damages.

> In this world too, the punishment was, ballpark, 100x direct damages. Which ballpark are you playing in to come to a number like that? I have friends whom are authors, and they've certainly not seen a penny from Anthropic. I somehow doubt they're the only ones out there that haven't been compensated for having their work blatantly stolen. In a just world we'd be ensuring Anthropic was destroyed as a result of their…

In a just world, copyright should be abolished with extreme prejudice should it's continued existence prevent humanity from developing aligned artificial super intelligence.

Re: The Kimi K3 Moment

#609

Earlier quoted context omitted.

So if you were to imagine the second largest economy in the world, governed by a monolithic politburo, with a long history of barely concealed programs of corporate and academic espionage, what do you think their approach would be to American technology worth trillions? Now take it a step further, this hypothetical government sees itself as the civilization representation of an ethnic group, and regularly attempts to…

The ego of America to think it is the only place where innovation can happen is pretty funny. Maybe America’s hubris and hostility towards talented foreigners is driving talent towards other countries? China doesn’t need to embed bad actors and spies in Anthropic and OpenAI, the U.S. is sabotaging itself just fine. America was at the forefront of technology because it attracted the best talent from all over the world…

If you think the number of US AI lab employees feeding information back to the homeland is zero, you’re terribly naive. There is the kind of money on the table people will absolutely lie, steal, and kill for.

There is also not strong evidence that immigration is the source of progress. We issued many more patents in the 19th century than now which was well before Hart-Cellar. The statement that immigration is even declining is not also supported by evidence.

Re: The Kimi K3 Moment

#610

Earlier quoted context omitted.

I could perhaps get myself to care just the tiniest bit if the information that was supposedly stolen wasn't generated by "stealing" from everybody else. Either it is fair use to train AI models on whatever information you can get your hands on for everyone or for no one.

Training data that’s still sitting in the pages of a book is not really all that useful.

that doesn't mean you get to steal it, and it means you don't get to complain when people reuse/steal it.
Post reply on HN