Live data from Hacker News

The Kimi K3 Moment

stephen.bochinski.dev

291–300 of 644 posts

Re: The Kimi K3 Moment

#291

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

assume you are a "second class lab" and you are in fact making progress by distilling the results of the frontier labs' efforts. what is the end game for this strategy? if the frontier labs shut down, or stop releasing to the public, and there's noting left to distill, how will you progress?

Distillation from a teacher model solves the self-start problem, that is, building a model to the point where it reason coherently. Without distillation, solving self-start is incredibly difficult since it requires millions of high quality training samples. Creating that kind of dataset takes an enormous amount of effort.

Once a model becomes competent enough to perform complex reasoning, a teacher model is no longer necessary. The model can now reason about its own behavior and build a better version of itself through recursive self-improvement (RSI).

Kimi K3 is capable of RSI.

Re: The Kimi K3 Moment

#292

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

Almost all markets depend on some form of regulation whether its as simple as "leave everyone alone but no stealing" or "every participant has to source every object through mountains of red tape." Thus far the US has not really chosen to go the Chinese rare-earth method yet. The problem with distillation attacks is the end result is everyone who is not doing them is going to deal with some kind of regulation whether…

    >  The problem with distillation attacks
I think it's worth stepping back here and pointing out the obvious. Y'all waging war on math. And I'm sorry, but that's the computing equivalent of legislating gravity.

Apologies for repeating myself here, but what you call "distillation" is function approximation.

I feel for the teams at Anthropic and Open AI, but unlike startups from prior eras; Anthropic and OpenAI have decided to be in the business of selling compute. Not creating a product that uses compute, but a product that's math running on compute. This is different from what Google is (or, rather was. As always, RIP Google 1998-2019).

Google's algorithm might be math, but Google search isn't. Google search is a process that's continuously operating in the background. Google crawls pages. Google stores and indexes what it finds. Google then exposes this to retrieval via its algorithm. User uses algorithm.

Now, let's compare that to AI models. When Anthropic serves Mythos / Opus etc, they're taking input or x from their user, doing compute, and then serving the result of the Mythos / Opus function, i.e.,

    f(text) -> (text_transform)
Where f is a continuous function, https://www.turing.ac.uk/sites/default/files/2025-11/languag...

According to Stone-Weierstrass, given enough values of y for f(x), anyone can approximate this function.

The fidelity and sophistication of this approximation definitely requires a lot of cleverness and effort, and it is arguably an imposition on Anthropic and OpenAI. But on a long-enough timeline, they don't even have to poll Anthropic or OpenAI. As the internet is flooded by PRs, content, emails written by Mythos / Claude, and just people otherwise sharing the results of Claude prompts, then there's an ever increasing set of data to approximate the f(x) that's f_Claude.

Eventually, in the future, anyone will be able to create a good enough approximation of the f_Mythos. Which is Anthropic's product.

Anthropic and OpenAI can now wage war on mathematics and the open-ended compute. Or, they can adapt and build a better product.

Choosing Option B was the Silicon Valley option / choice. I think the OG large-scale Valley lobbying effort, the Semiconductor Industry Association, was unique in that it prioritized and chose to do real research.

https://en.wikipedia.org/wiki/Semiconductor_Industry_Associa...

https://en.wikipedia.org/wiki/Semiconductor_Research_Corpora...

This helped the industry to survive and outcompete the pressure they were facing (at the time).

Re: The Kimi K3 Moment

#293
post #258
post #8

I tried Kimi K3 on a task I've done with every other model I use regularly ( https://swelljoe.com/post/i-let-every-agent-implement-its-ow... ) and found it chewed a lot longer on the problem and ate up almost the entirety of a 5 hour usage limit on their $19 plan. I only have the $20 plan from OpenAI and the same task, with a lot of the same implementation details as Kimi Code, only took a few minutes and consumed al…

We really need to stop using $/M tokens as the pricing benchmark. I've found that the number of tokens used tends to be a bigger factor than the listed per token price. The cost per task vs. intelligence curve is really what you care about, and in my estimation Chinese models are just not there. They are focused on benchmaxing and getting the highest raw score they can, rather than efficiency.

The artificialanalysis cost per task chart has DeepSeek as the clear winner and Fable as the clear loser. But I would still pick Fable for some tasks, so that also can't be all there is to it.

But I agree that price per token figure is not great. It seems even the tokens per character can vary between models, so it's basically useless.

Re: The Kimi K3 Moment

#294

Earlier quoted context omitted.

Anthropic paid $1.5 billion in compensation to copyright holders for use of their content in training data. The payment was for illegally downloading copyrighted material, not training. Training was explicitly ruled to be fair use.

Partially correct. The court explicitly ruled that training on pirated data, which is what Anthropic was doing, is not considered fair use. Training on legally acquired / licensed data is potentially fair use.

It's not potentially, it's settled. At least for now as neither case wanted to move on to appeals

Re: The Kimi K3 Moment

#295

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

The fact that API based distillation is even a conversation right now makes me feel like the U.S. has their heads so far in the sand that it’s not really excusable. These Chinese labs are producing novel models, publishing their techniques and sharing their open weights and the first topic of conversation is how they stole from U.S. AI labs. Setting aside the fact that it doesn’t make any feasible sense to do API dis…

> We have to stop crying distillation, it’s getting embarrassing and at this point feels even a bit delusional.

It's a PR campaign - when they say its an "attack" they don't mean on Anthropic - but on America itself. What kind of American can let such a brazen attack go unanswered? At the very least, they ought to demand the dangerous, pinko, stolen models be banned in all 50 states, and pay whatever price demanded by the patriotic, freedom-loving, all-American AI labs that can never be accused of stealing.

Re: The Kimi K3 Moment

#296
post #53

Earlier quoted context omitted.

Doesn't read like AI writing whatsoever.

I don't know. You can't just rely on looking for em dashes or other obvious tells because anyone who cares can get the AI to avoid those. > When the headline model on your plan can be switched off because the economics don’t work, the plan was never really selling you the headline model. Kimi’s tiers don’t come with that asterisk. This line has a certain smug, punchy cleverness that I associate with AI. To me, the vi…

> You can't just rely on looking for em dashes or other obvious tells

I didn't.

Re: The Kimi K3 Moment

#297

Earlier quoted context omitted.

The fact that API based distillation is even a conversation right now makes me feel like the U.S. has their heads so far in the sand that it’s not really excusable. These Chinese labs are producing novel models, publishing their techniques and sharing their open weights and the first topic of conversation is how they stole from U.S. AI labs. Setting aside the fact that it doesn’t make any feasible sense to do API dis…

There's little doubt that Kimi K3 was distilled off Claude. Anthropic stated in February that Moonshot AI (the creator of Kimi) distilled ~3.4 million exchanges from Claude models, as explained in their press release https://www.anthropic.com/news/detecting-and-preventing-dist...

While it sounds like a lot, do you suppose 3.4 million sessions come even close to being sufficient to train a frontier model?

Assuming each session was 10,000 words each, that's 34 billion words; lets call it 50 billion tokens (0.05 trillion) unfairly pilfered from Claude. That left Moonshot needing to scrounge for the other 14.950 trillion training tokens required for a baseline frontier model.

Re: The Kimi K3 Moment

#298
post #177
post #22

Earlier quoted context omitted.

Is Kimi K3 subsidized as hard as the other models out there?

From the end of the month it will be served profitably by other providers around the world, like Kimi K2.6

Sure, but at those API rates it doesn't beat the subscription plans for OpenAI.

Re: The Kimi K3 Moment

#299

Earlier quoted context omitted.

Yes yes, we all understand the game-theoretic race-to-the-bottom you're describing here. Somehow despite linux being FOSS it still powers most of the important computing in the world. Can you explain how that works despite it being free? Once you understand that case I think you'll understand the game-theory behind how large projects can exist in the absence of traditional IP protection.

the obvious difference is the massive scale of data and compute required to develop and evolve these models, and the costs they impose on those building them.

Smaller budgets, slower improvement, less risk. They're not entitled to profits if that business model isn't sustainable. They're not entitled to a change in IP laws to protect their business model. They're not entitled to growing that fast.

Re: The Kimi K3 Moment

#300

Earlier quoted context omitted.

Us models didnt pay for licenses too

That is incorrect. Anthropic paid $1.5 billion in compensation to copyright holders for use of their content in training data. OpenAI pays hundreds of millions per year across 150+ licensing deals for access to copyrighted data. Meta and Alphabet have similar arrangements. Under the settlement, Anthropic was forced to delete the pirated data they were training on. Chinese labs can still train on pirated data. I doubt…

They used two of my books and I'm still waiting for my cheque here.
Post reply on HN