Live data from Hacker News

The Kimi K3 Moment

stephen.bochinski.dev

241–250 of 644 posts

Re: The Kimi K3 Moment

#241

Earlier quoted context omitted.

> Distillation “attacks” are not attacks. If "distillation attacks" happen, we have to conclude there is some value add in what model labs do. Regardless of how we feel about using existing human knowledge in the way they currently do, it's simply impractical to infer that everything that happens downstream of LLMs can not be an attack on some IP because of it. So both things can be true: a) People infringe on Anthro…

Regardless of whether it’s intellectual property or it isn’t intellectual property, it doesn’t actually matter. If AI doesn’t stop seeing diminishing returns in scaling up, and it hasn’t yet in the 10 years since the attention/transformers paper, the advent of AI will be the most important development in the history of humanity. Controlling that machine, or at least having one of your own, is an existential problem f…

In fact, China stealing fire from the gods is essential to the future balance of power, so long as they keep making the results freely available.

Re: The Kimi K3 Moment

#242
post #232
post #8

I tried Kimi K3 on a task I've done with every other model I use regularly ( https://swelljoe.com/post/i-let-every-agent-implement-its-ow... ) and found it chewed a lot longer on the problem and ate up almost the entirety of a 5 hour usage limit on their $19 plan. I only have the $20 plan from OpenAI and the same task, with a lot of the same implementation details as Kimi Code, only took a few minutes and consumed al…

Aren't they still locking reasoning to "max" pending adjustments to support shorter reasoning levels.

It let's me choose different thinking levels in Kimi Code. Not sure if it actually works, yet, but it says "Thinking set to high." when I change it from max.

Re: The Kimi K3 Moment

#243

Earlier quoted context omitted.

> empirically, it appears that distillation of a more advanced model is a required first step I see no evidence for that. > if this were not the case, then we would be observing chinese models that far surpass frontier models It's pretty clear that the primary reason for the difference is budget and compute availability. Chinese labs have at least an order of magnitude less money than Anthropic and OpenAI. > what hap…

https://www.anthropic.com/news/detecting-and-preventing-dist... Moonshot AI Scale: Over 3.4 million exchanges The operation targeted: Agentic reasoning and tool use Coding and data analysis Computer-use agent development Computer vision Moonshot (Kimi models) employed hundreds of fraudulent accounts spanning multiple access pathways. Varied account types made the campaign harder to detect as a coordinated operation.…

>request metadata, which matched the public profiles of senior Moonshot staff

Translation: we have the machinery in place to identify our users, and actively do so.

Re: The Kimi K3 Moment

#244
post #51

Kimi K3 is really good, but it’s obviously worse than Fable, usually worse than Opus, in my experience.

That’s not obvious to me at all. Especially your claim around opus.

Agreed. I think benchmarks are pretty much right in pegging it somewhere in between Opus and Fable.

Re: The Kimi K3 Moment

#245
post #118
post #8

I tried Kimi K3 on a task I've done with every other model I use regularly ( https://swelljoe.com/post/i-let-every-agent-implement-its-ow... ) and found it chewed a lot longer on the problem and ate up almost the entirety of a 5 hour usage limit on their $19 plan. I only have the $20 plan from OpenAI and the same task, with a lot of the same implementation details as Kimi Code, only took a few minutes and consumed al…

Absolutely do not pay for the kimi plans thinking they will be cheaper. If you sign up with a Chinese phone number, you can get the same plan for 200 yuan instead of 200 usd, it also only accepts Chinese payment methods iirc. So the plans are really made for Chinese userbase.

That sounds complicated. I'll just use my month of Kimi and then cancel. I have too many AI subscriptions to use them all, anyway. I subscribed mostly to test it. I mean, if it turned out to be competitive, I would keep it, but if it doesn't turn out to really excel and anything and also take longer than Claude or OpenAI models, I'll stick with them.

Re: The Kimi K3 Moment

#246

Earlier quoted context omitted.

In my opinion, for the vast majority of use cases, DeepSeek is still the most cost-effective model by a mile. $10 feels like it lasts forever.

Yep, with Reasonix, DeepSeek is free real estate. Seems to just go and go for pennies. And, DeepSeek is what I use for any task that works best with an API. It's cheap enough to where I don't think about cost, made even cheaper by DeepSeek having the most effective and cheap caching in the industry, and it's good enough to where I rarely have to follow up with a more expensive model or manually fix things. It's been…

> It's cheap enough to where I don't think about cost, made even cheaper by DeepSeek having the most effective and cheap caching in the industry

I use DeepSeek v4 Flash & MiMo v2.5 Pro. Prefer the latter over DeepSeek v4 Pro because it costs the same while being equally good & less chattier for coding workloads. Although, I've begun experimenting with Hy3 (as an in-between Flash & Pro) & GLM 5.2 (for long-horizon tasks).

Re: The Kimi K3 Moment

#247
post #206

Earlier quoted context omitted.

> Subscription usage limits are hard to measure as none of the providers tell you directly what it means in terms of tokens or anything else you can easily compare AI subscription pricing is so goofy. You get some amount of usage that varies by models, is measured by opaque token usage, driven by how many tokens the (usually) vendor-provided interface (or model itself) wants to use. Then your usage is limited by time…

AI subscription pricing was fine when it was $100/month for some opaque 5 hour token budget I don't think I ever used, not even that one day where I coded for 14 hours non-stop using Fable. But like most people with low token usage, I had a human in the loop and and I didn't use workflows with swarms of agents. Now, of course, the plan is to remove Fable from the subscription. To paraphrase Darth Vader, they have alt…

Gpt 5.6 is still like this at least for the $200/month option. It’s also always faster than fabel. Fabel might be able to do some things better but I don’t have time to constantly wait and find out.

Re: The Kimi K3 Moment

#248

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

The fact that API based distillation is even a conversation right now makes me feel like the U.S. has their heads so far in the sand that it’s not really excusable.

These Chinese labs are producing novel models, publishing their techniques and sharing their open weights and the first topic of conversation is how they stole from U.S. AI labs.

Setting aside the fact that it doesn’t make any feasible sense to do API distillation, these models are outperforming frontier models on a number of benchmarks, and often times run more efficiently by several orders of magnitude.

We have to stop crying distillation, it’s getting embarrassing and at this point feels even a bit delusional.

Re: The Kimi K3 Moment

#249

Earlier quoted context omitted.

> People infringe on Anthropics IP Unless someone literally stole the weights somehow (which is not out of the question, I doubt either oAI/Anthropic have the capabilities to prevent a state-level actor getting those weights), distillation from generations is not infringement on anyone's IP nor is it stealing nor is it an attack. It can't be. As long as you pay for tokens you get to do whatever you want with them. So…

Its definitely an attack. Thats established from anthropics perspective. No one has a right to use Anthropic’s services in ways that directly violate the ToS and user agreements.

> Its definitely an attack. Thats established from anthropics perspective.

How do things get "established" from someone's perspective, exactly?

By that logic it is established from my perspective that Anthropic has no right to train on anything I've written that is publicly available on the internet.

Of course, they don't care about my perspective, but then again I don't care about theirs.

Re: The Kimi K3 Moment

#250
post #24

The current administration's immigration policy isn't helping. This wouldn't have happened 10 years ago because the US was this city on the hill that everyone wanted to immigrate to. Talented Asian researchers would have immigrated to the US and China would be deprived of talent.

Your comment feels like an outdated brain drain model where talented Chinese researchers naturally want to leave China and the only question is whether the US lets them in. That may have been closer to reality 10-20 years ago, China is a different country now, what I mean by that is they offer research funding, they have huge digital behemoths (alibaba, tencent, huawei, bytedance etc), large scale deployment opportun…

Chinese students still want to attend US universities [1]. While it is true that the progress made by China is a factor, this administration's policies are the bigger deterrent [2] [3].

[1] https://www.latimes.com/world-nation/story/2025-02-21/why-ch...

[2] https://www.wsj.com/world/china/americas-allure-fades-in-chi...

[3] https://www.theguardian.com/world/2025/jun/06/chinese-studen...

Post reply on HN