Live data from Hacker News

The Kimi K3 Moment

stephen.bochinski.dev

421–430 of 644 posts

Re: The Kimi K3 Moment

#421

Earlier quoted context omitted.

Shrug. Hard to feel sorry for companies that created their empires by ignoring copyright themselves. Also, 'most extensive industrial espionage campaign, probably ever' is absolute nonsense. They did not need to infiltrate the companies for this nor are you accusing them of stealing any trade secrets. This is only about whether they looked at their competitors' products from the outside (in the form of conversation t…

> 'most extensive industrial espionage campaign, probably ever' is absolute nonsense You are completely underestimating the scale of what is happening here. Chinese AI labs are actively facilitating an industrial-scale network of tens of thousands of bot accounts, that resell Claude tokens at 97% below official API prices. They buy subsidized Max 5x plans (sometimes with stolen credit cards), then split the subscript…

Does anything about that strike you as particularly unfair, given the moral compass defined by Big AI?

Re: The Kimi K3 Moment

#422

I think it's the opposite. Kimi K3 has 2.8 trillion parameters. We don't know the number of parameters of ChatGPT 5.6 or Opus 4.8, but it's probably in the same region. Fable/Mythos are rumored to be around 10 trillion. So, K3 is directly comparable with ChatGPT 5.6 and Opus 4.8, and the price is not so much lower: K3: $3/$15 per 1 Mtok input/output ChatGPT 5.6 Sol: $5/$30 Opus 4.8: $5/$25 This is not a watershed mom…

And how much token and time you use to solve a problem? The price alone doesn't mean anything.

Example DeepSeek-V4-Pro (high) needs 10 times more token then GPT 5.5 (medium) and can compete only with the price.

The real price saver are the cache prices, the ting, that nearly nobody has on their radar.

Re: The Kimi K3 Moment

#423
post #337

Earlier quoted context omitted.

Free server racks for everyone when the bubble bursts!

"Free server racks for everyone when the bubble bursts!" I actually have this trophy from the previous bursting bubble ... a Sun microsystems rack populated with three e4500. $750k + of equipment at original list price ...

> Sun microsystems rack populated with three e4500

That's super cool; I bet that's a lot of fun to play around with. I wonder how much of this stuff just ended up in a landfill because it was too much effort to find buyers.

Re: The Kimi K3 Moment

#424

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

[dead]

Re: The Kimi K3 Moment

#425
post #73

Well, there is the small issue of privacy policy: Kimi will train their models on your interactions if you use their subscriptions, and only with direct API usage (billed at API prices) they say they won't. Whether you trust that is another matter. Those things do make a difference to some of us, even though nothing is black and white. In my case, I'll probably want to wait until other providers appear through OpenRo…

> In my case, I'll probably want to wait until other providers appear through OpenRouter and then I'll try to judge how much I trust them.

Keep in mind that the Moonshot team have identified multiple providers who configure their setup wrong which make their model perform worse than expected. This is why Moonshot created Kimi Verifier, but I guess its up to the provider if they want to do that.

Re: The Kimi K3 Moment

#426

Earlier quoted context omitted.

Kimi calling itself claude means nothing. During pre-training, when the model learns to "simulate" the internet text, it will naturally be fed with a bunch of data about Claude and ChatGPT. With the amount of LLM outputs on the internet today, it is not surprising at all that a model would naturally call itself Claude or ChatGPT. You can mitigate that in post-training (or actually in pre-training as well) by training…

Sure, but then Qwen should leak that too, and it doesn't. K3 calls itself Claude 7 out of 48 times, Qwen does it 0 out of 48, and the only other model to identify itself as Claude is DeepSeek. and DeepSeek is alleged to also distill from Claude data anyway. So this isn't something every model absorbed from the same web text. And you skipped over the strongest datapoint that K3 is distilled: K3 reproduces Claude's pub…

> Qwen should leak that too, and it doesn't.

FWIW I had Qwen identifying itself as "a language model made by Google" in one conversation, although I could not reproduce this reliably.

Re: The Kimi K3 Moment

#427

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

The desire to accuse China of just copying is like 20 years out of date. It’s been wrong since some people on HN were in diapers. People are going to be gobsmacked when, in our lifetime, China becomes a world power comparable to the U.S. Probably still poorer per capita, but at Spain/Italy levels, not third world country levels. And they’ll be shocked at the implications of that on the world economy, migration patter…

[dead]

Re: The Kimi K3 Moment

#428
post #414

Earlier quoted context omitted.

There are huge evidence of copying. Some day China can pioneer in science or technology but the current claim about Chinese companies leading AI development is ridiculous given the evidence of distillation and the fact that like 95 percent of science that lead to the current state of AI happened in either North America or Europe. To be honest if you want to list academic papers that lead to the current AI models the…

In 2017 maybe. This chart shows last year’s Neurips accepted papers by country and institution (top 50). What is missing here is that the papers from American institutions also have mostly Chinese authors. Europe is sliding and Singapore has more papers than Canada. There is a clear trend. https://www.reddit.com/r/accelerate/comments/1pi64q0/papers_... US is still winning because of their hardware dominance. Also the…

Volume is up because AI generation help writing papers. We should find a better measure of impact. I like to see the same charts on best paper awards.

Re: The Kimi K3 Moment

#429
post #258

Earlier quoted context omitted.

We really need to stop using $/M tokens as the pricing benchmark. I've found that the number of tokens used tends to be a bigger factor than the listed per token price. The cost per task vs. intelligence curve is really what you care about, and in my estimation Chinese models are just not there. They are focused on benchmaxing and getting the highest raw score they can, rather than efficiency.

The artificialanalysis cost per task chart has DeepSeek as the clear winner and Fable as the clear loser. But I would still pick Fable for some tasks, so that also can't be all there is to it. But I agree that price per token figure is not great. It seems even the tokens per character can vary between models, so it's basically useless.

Wow, you weren't kidding. I looked at their chart, and the cost-per-task for Fable is more than double Sol's. And DeepSeek absolutely stomps. Four cents per-task vs Sol's $1 and Fable's $3.

I might need to check out DeepSeek more. I had no idea the difference was this obscene. Makes me wonder if something's off with the benchmark. A 70x cost reduction vs. Fable seems too good to be true.

Re: The Kimi K3 Moment

#430

Earlier quoted context omitted.

Shrug. Hard to feel sorry for companies that created their empires by ignoring copyright themselves. Also, 'most extensive industrial espionage campaign, probably ever' is absolute nonsense. They did not need to infiltrate the companies for this nor are you accusing them of stealing any trade secrets. This is only about whether they looked at their competitors' products from the outside (in the form of conversation t…

> 'most extensive industrial espionage campaign, probably ever' is absolute nonsense You are completely underestimating the scale of what is happening here. Chinese AI labs are actively facilitating an industrial-scale network of tens of thousands of bot accounts, that resell Claude tokens at 97% below official API prices. They buy subsidized Max 5x plans (sometimes with stolen credit cards), then split the subscript…

if Ford bought hundreds of millions of dollars worth of Hyundais, put extra instrumentation in them, and resold them at a discount to customers who agreed to the instrumentation in exchange for the discount, is Ford doing industrial espionage?
Post reply on HN