Live data from Hacker News

The Kimi K3 Moment

stephen.bochinski.dev

311–320 of 644 posts

Re: The Kimi K3 Moment

#311
One thing is absolutely clear - open-source models have already reached the level of top commercial ones. The last bottleneck is hardware, but that threshold is decrasing fast too. While it is still extremely expensive to run Kimi K3 model at home, there are already many very capable free models you can run on decent hardware. This trend will definitely continue.

Re: The Kimi K3 Moment

#312
Since Kimi’s paid plans are mentioned in the article..interested ones should know that you can only access 1M context model with $79/mo or higher plan; otherwise you are capped at 256k context. Also, with minimal $15/mo plan k3 is currently not supported at all. (prices are yearly plan discount prices)

ref: https://www.kimi.com/code/docs/en/kimi-code/models.html

Re: The Kimi K3 Moment

#313

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

> There was never any plausible explanation for why this wouldn’t happen.

What a nice post hoc revision of history. Distillation is still an active area of research, that you can distill models as easily as you can it genuinely interesting and absolutely not something that was taken for granted even 12 months ago.

Even 6 months ago this idea that 'using model outputs as training examples' was listed as the reason that all models would fail in the near future due to some spooky circular training catastrophe.

Don't pretend like this was so obvious.

Re: The Kimi K3 Moment

#314
post #312

Since Kimi’s paid plans are mentioned in the article..interested ones should know that you can only access 1M context model with $79/mo or higher plan; otherwise you are capped at 256k context. Also, with minimal $15/mo plan k3 is currently not supported at all. (prices are yearly plan discount prices) ref: https://www.kimi.com/code/docs/en/kimi-code/models.html

Thanks for mentioning that. I also wanted to use API only and with the cache hit rates I'm getting with Reasonix/whale harness on deepseek, it's going to be a difficult adjustment moving away from practically free.

Re: The Kimi K3 Moment

#315

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

> There was never any plausible explanation for why this wouldn’t happen. What a nice post hoc revision of history. Distillation is still an active area of research, that you can distill models as easily as you can it genuinely interesting and absolutely not something that was taken for granted even 12 months ago. Even 6 months ago this idea that 'using model outputs as training examples' was listed as the reason tha…

I think you’re being overly combative. It’s intuitively quite obvious that it’s incredibly easy to implement and the circular training catastrophe was only ever a conjecture. It’s kind of like releasing a crypto primitive without knowing a proof. Like… maybe it works, but you can’t assume that just because you don’t know how to break it. You have to remember that 100s of billions of enterprise valuation rely on frontier models being moats. The burden of proof is on those raising valuations assuming they will capture the full market.

Re: The Kimi K3 Moment

#316

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

> There was never any plausible explanation for why this wouldn’t happen. What a nice post hoc revision of history. Distillation is still an active area of research, that you can distill models as easily as you can it genuinely interesting and absolutely not something that was taken for granted even 12 months ago. Even 6 months ago this idea that 'using model outputs as training examples' was listed as the reason tha…

I agree that hindsight is doing work here, but DeepSeek R1 from Jan 2025 seemed to heavily leverage distillation, and 18 months is an eternity in this climate.

Re: The Kimi K3 Moment

#317

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

I suspect that distillation attacks may be slightly exaggerated. Most of the training data used during fine-tuning is now synthetic data. You can't just repeat the same stuff twice, therefore another LLM is writing a text book that is explaining a topic in detail, ideally without any gaps in the material.

Re: The Kimi K3 Moment

#318
post #115

Terms of use are very broad and not friendly for most things. Can’t use for commercial purposes. Can’t opt out of training. Data retained.

Correct: can't opt out of training. This is well documented. "Can't use for commercial purposes" - incorrect AFAICT. In what sense do you mean this? The open weight MIT version obviously allows for commercial use, but I don't think that's what you're referring to, because training data is irrelevant on the open weight version. Pretty sure the API allows commercial use too. Maybe the free version doesn't? But who care…

> Service Misuse. You acknowledge that without the written consent of us and/or the relevant rights holders, (i)you have no authority to use Kimi and the content generated by Kimi in any commercial manner; (ii)you may not use our Services to develop products or services that compete with us.

https://www.kimi.com/user/agreement/modelUse

Re: The Kimi K3 Moment

#319

Earlier quoted context omitted.

>People infringe on Anthropics IP No. Authors do not infringe on IP when they read another's book, nor should the lumber company be able to dictate how I use planks and if I can resell them if i'm done with them. You're framing it as if the added value of the author or lumber company, awards them consideration when somebody uses the products to create more value. IP law was always a big mess, and these questions cros…

It's more simple: They infringe on the IP by way of violating the ToS. If you violate ToS and the company suffers financial harm, they usually can (usually) sue you in civil court for damages.

Terms of service are separate from intellectual property

Re: The Kimi K3 Moment

#320

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

Well, there is precedence: Google can scrape the web, but you can't scrape Google. Laws around compiled databases exist for a reason: you can't just copy the phone book if effort has gone into compiling it, it is itself copyrightable

This is the opposite of legal reality, at least as far as the US is concerned.
Post reply on HN