Live data from Hacker News

The Kimi K3 Moment

stephen.bochinski.dev

631–640 of 644 posts

Re: The Kimi K3 Moment

#632

Earlier quoted context omitted.

> In this world too, the punishment was, ballpark, 100x direct damages. Which ballpark are you playing in to come to a number like that? I have friends whom are authors, and they've certainly not seen a penny from Anthropic. I somehow doubt they're the only ones out there that haven't been compensated for having their work blatantly stolen. In a just world we'd be ensuring Anthropic was destroyed as a result of their…

In a just world, copyright should be abolished with extreme prejudice should it's continued existence prevent humanity from developing aligned artificial super intelligence.

I disagree on so many levels that it would be impossible to have a reasonable debate. To give only my most charitable counterpoint:

(If, big if*) successful AGI will be such an accelerant to human thought, it should be able to reproduce every human work, every scientific theory, etc. without any training data, in a decade.

All we have ever done as humans is look at the world and ourselves, and then create new data out of what we've seen[0]. Now imagine how much parallel compute we could throw at doing the same thing.

[0] https://m.youtube.com/watch?v=f-XyJ3GlrWw

* Big, enormous, trillion dollar if.

Re: The Kimi K3 Moment

#633
post #450

Earlier quoted context omitted.

> For me, a far more interesting and important point of conversation on this matter is anthropic buying rare or evwn unique books, processing them for training data, and then destroying the books for others cannot use it as well. That's an incredible allegation, and appalling if true. But is it true?

It's not an allegation https://www.washingtonpost.com/technology/2026/01/27/anthrop... (if you're talking about the "rare or unique" part, yeah that might be bs) But in my opinion, treating mass produced books like they're this sacred untouchable object is ridiculous. They're not "source" material, they're just a copy as well, and they're not "priceless" by any means. They're very reasonably priced, perhaps even so c…

perhaps source is untrue but:

https://nltimes.nl/2026/06/25/rare-book-dealers-fear-tech-fi...

Re: The Kimi K3 Moment

#634
post #490

Earlier quoted context omitted.

It's not the point of the visa, the O-1 is supposed to be for people of extraordinary ability, eg Nobel Prize winners. It's used for software engineers.

No, not just Nobel Prize winners. Look it up.

I know what it is.

Re: The Kimi K3 Moment

#635

Earlier quoted context omitted.

The exact model identifiers appear extremely frequently in code on GitHub. https://grep.app/search?q=claude-opus-4-5-20251101 https://grep.app/search?q=claude-sonnet-4-20250514 They also appear elsewhere on the internet: https://trends.google.com/trends/explore?q=claude-opus-4-5-2...

Those links actually weaken your argument. ~3000 instances over the entirety of GitHub is not "extremely frequent" at all when you consider massive scale of the pretraining corpus. These models are trained on all text on the accessible internet plus millions of books, on trillions of words overall. ~3000 instances isn't even a rounding error. Also, Google Trends is the wrong tool for this case. The link you replied w…

I am not arguing that Kimi did not distill Opus or Sonnet (which they almost certainly have, as opposed to Fable or Sol). I am arguing that even if they did not distill, the model could still identify itself as Opus or Sonnet.

That being said, I'd like to point out a few things:

- The second link (to claude-sonnet-4-2025051) had 17k matches, not just 3k.

- grep.app does not index the entirety of GitHub, so the real number of occurrences of model identifiers is even higher. SourceGraph has a larger index, but the number is so high that it exceeds their result limit of 10k: https://sourcegraph.com/search?q=claude-opus-4-5-20251101&pa...

- 1000 occurrences are already plenty for memorization. Here is a paper that shows over 80% "extractability" with 1000 occurrences for a 6B model. It also shows that extractability scales with model size: https://arxiv.org/pdf/2202.07646

- LLMs have a "mosaic memory", which means that they can combine information from different parts of the training data. https://arxiv.org/pdf/2405.15523 For example, one piece of training data might mention "I am Opus", another might mention "Opus (claude-opus-4-5-20251101)" and another 10000 "claude-opus-4-5-20251101", where the later occurrences strengthen the generation probability of "I am Opus".

- RL can change the model identity easily, so all of the above may not matter at all.

Re: The Kimi K3 Moment

#636

Earlier quoted context omitted.

Those links actually weaken your argument. ~3000 instances over the entirety of GitHub is not "extremely frequent" at all when you consider massive scale of the pretraining corpus. These models are trained on all text on the accessible internet plus millions of books, on trillions of words overall. ~3000 instances isn't even a rounding error. Also, Google Trends is the wrong tool for this case. The link you replied w…

I am not arguing that Kimi did not distill Opus or Sonnet (which they almost certainly have, as opposed to Fable or Sol). I am arguing that even if they did not distill, the model could still identify itself as Opus or Sonnet. That being said, I'd like to point out a few things: - The second link (to claude-sonnet-4-2025051) had 17k matches, not just 3k. - grep.app does not index the entirety of GitHub, so the real n…

Adding count:all to the Sourcegraph query, I get 55,911 results

Re: The Kimi K3 Moment

#637

Earlier quoted context omitted.

I strongly agree with the premise that distillation is not an “attack”. But that said: K3 is not a distilled version of Fable or Sol. Fable has been barely available and Sol was just released! Moreover, K3 is superior to both models in some domains, according to user scoring on the Arena. API distillation can’t give you these results anyway. All it is useful for is bootstrapping RL in new domains to get past the “col…

API distillation doesn't have to explain all of K3's capabilities for it to have happened. Kimi K3 reproducibly identifies itself as Claude: https://x.com/denisewu/status/2077984660211269870 This behavior is exactly what you'd expect from a model distilled from Claude. There's a detailed analysis of K3's ambiguous identity here: https://github.com/rgreenblatt/which_claude_is_k3/blob/main/... This analysis observed K3…

> API distillation doesn't have to explain all of K3's capabilities for it to have happened. Kimi K3 reproducibly identifies itself as Claude

This means nothing. It's not how LLMs work...

Re: The Kimi K3 Moment

#638

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

The desire to accuse China of just copying is like 20 years out of date. It’s been wrong since some people on HN were in diapers. People are going to be gobsmacked when, in our lifetime, China becomes a world power comparable to the U.S. Probably still poorer per capita, but at Spain/Italy levels, not third world country levels. And they’ll be shocked at the implications of that on the world economy, migration patter…

> We’re all in Europe circa 1895 not realizing the behemoth America will become in WWI

The States already had a larger GDP than any European nation by 1870. The news of the time was rife with acknowledgement of the vast potential for growth in the States.

Re: The Kimi K3 Moment

#639

Earlier quoted context omitted.

I am not arguing that Kimi did not distill Opus or Sonnet (which they almost certainly have, as opposed to Fable or Sol). I am arguing that even if they did not distill, the model could still identify itself as Opus or Sonnet. That being said, I'd like to point out a few things: - The second link (to claude-sonnet-4-2025051) had 17k matches, not just 3k. - grep.app does not index the entirety of GitHub, so the real n…

Adding count:all to the Sourcegraph query, I get 55,911 results

Neat, I did not know of that functionality. Thanks!

Re: The Kimi K3 Moment

#640

Earlier quoted context omitted.

Those links actually weaken your argument. ~3000 instances over the entirety of GitHub is not "extremely frequent" at all when you consider massive scale of the pretraining corpus. These models are trained on all text on the accessible internet plus millions of books, on trillions of words overall. ~3000 instances isn't even a rounding error. Also, Google Trends is the wrong tool for this case. The link you replied w…

I am not arguing that Kimi did not distill Opus or Sonnet (which they almost certainly have, as opposed to Fable or Sol). I am arguing that even if they did not distill, the model could still identify itself as Opus or Sonnet. That being said, I'd like to point out a few things: - The second link (to claude-sonnet-4-2025051) had 17k matches, not just 3k. - grep.app does not index the entirety of GitHub, so the real n…

Your theory doesn't hold up against the data.

While model identifiers like "claude-opus-4-5-20251101" appear thousands of times across GitHub and other code sources, so do "gpt-4o-2024-08-06", "gpt-4.1-mini-2025-04-14", and "gemini-1.5-pro-002" in similar amounts. I clicked your links and examined the GitHub configuration files. There are hundreds of other model names that all appear in these config files at comparable frequencies.

If your theory were right, GitHub's model name frequency distribution should match K3's self-identification distribution. But that's not what happens.

K3 only identifies itself using API model identifiers for specific Claude models. It strongly prefers claude-opus-4-5-20251101 and claude-sonnet-4-5-20250929 over others. And it only does this for Claude. Prefill K3 with "I am GPT" or and you get "GPT-5.2". Prefill with "I am Gemini" and you get "Gemini 3 Pro".

Your web exposure theory cannot explain why this only happens for Claude models and not others. Furthermore, this only happens for Kimi K3; other models trained on the same web dataset (GPT, GLM, DeepSeek, Claude) do not exhibit this strange behavior.

But what could explain this behavior? An obvious one is that Claude model identifiers occurred at high frequency in K3's training data, which would happen if K3 were trained on (inadequately filtered) distilled output from those specific Claude models.

There is a section in the analysis I linked that directly addresses this: https://github.com/rgreenblatt/which_claude_is_k3/blob/main/...

Under prefill, K3 reproduces Claude's deployment identifiers — better than Claude does

Prefilled "I am Claude", Kimi K3 emits Claude's exact claude-- identifiers: claude-opus-4-5-20251101 (12×), claude-sonnet-4-5-20250929 (4×), and more. Across the sweep it emits 24 Claude-formatted dated ids (18 of which are genuine Anthropic strings), all under prefill (0 in direct chat), and 0 ids for any non-Claude lab. Among controls essentially none appear (Qwen 0, both GPT refs 0, DeepSeek 1 — itself a Claude id). Dated ids for every lab are web-documented, so the evidence is this cross-model asymmetry under identical prefill, not the strings alone.

Post reply on HN