Live data from Hacker News

Kimi K3 Architecture Overview and Notes

sebastianraschka.com

71–80 of 125 posts

Re: Kimi K3 Architecture Overview and Notes

#71
post #60

Earlier quoted context omitted.

This is so true honestly.

Also Reddit-style humorous replies were frowned upon because they ruin serious discussion and encourage karma chasing as opposed to providing valuable insight.

To be honest, comment like yours were also frowned upon haha. If you think something is bad for discussion, downvote. If you think something is objectionable, flag it. And now I have made a faux pas by explaining the rule to you rather than just downvoting and moving on. Its turtles all the way down today.

Re: Kimi K3 Architecture Overview and Notes

#72
post #65

Earlier quoted context omitted.

"Kimi is largely a byproduct of distillation" and "Kimi is introducing new and novel approaches" are not mutually exclusive, and I'm not sure it's clear from the paper how much of the improvement comes from the new approaches. So I wouldn't take the new approaches to be much evidence about whether the distillation attacks occurred.

Even Anthropic hasn't claimed "Kimi is largely a byproduct of distillation" There's been some distillation. Just like how Elon said in court they distill to make Grok.

I thought there is discussion circulating around whether the main reason Kimi is impressive is due to distillation. Though possibly if this occurred, this was just one component of their training pipeline and not a majority

Re: Kimi K3 Architecture Overview and Notes

#73
post #67

Genuine question: how reproducible / usable / verifiable are these architectures from the published documentation? Are they similar to PDF/DWG/PSD specifications, where the format look like an open spec at first sight until you attempt to implement it and realize the crucial implementation details are undocumented?

It's entirely reproducible from the available documentation (which is why you see vLLM, SGLang, MLX etc all racing to produce optimized implementations). (As an aside, this is why the "open weights are not open source" thing is a complete misunderstanding. The weights themselves along with the documentation give you enough to fine tune the LLM. You can't rebuild it from scratch, but you can't do this even with the da…

I agree that if you have the weights you can use/train a model with the same architecture, and that you won't get the exact weights on your own due to randomness. But isn't data an extremely important part of your ability to effectively train/finetune? It might be much harder to get close to the level of the open weight model if you don't have the data that made it, which is why I think the open weights vs open source distinction is useful.

Re: Kimi K3 Architecture Overview and Notes

#74

So, unlike what leaders of western labs labs would like you to believe (that Kimi is just the result of distillation attacks), they are introducing new and novel approaches.

Moreover, Distillation is a misused term here. Distillation means training a smaller student model using a larger teacher model to completely mimic its behavior. As in you take a base model and create its smaller “turbo” version. Like distillation in Chemistry it means it’s %100 purified version of its teacher.

Re: Kimi K3 Architecture Overview and Notes

#75
post #61

Anybody getting the result that Kimi 3 is more expensive than Opus 5 or Sol on Cursor? Pretty sure Kimi 3 sucked up a good chunk of my ultimate plan in a few prompts. Anyone have any tools or ways to understand per model usage towards cursor subscriptions? I know there are alternatives to cursor just haven’t made the move yet. (Edit spelling)

The assertions doing the rounds that Kimi K3 and GLM 5.2 are way cheaper than Claude/GPT are not true -- DeepSeek V4 Pro is a lot cheaper but K3 and 5.2 ain't. Turns out that you actually have to fork out some cash for frontier-esque models, be they Chinese or American. Hope that helps.

Source: my bank balance

Re: Kimi K3 Architecture Overview and Notes

#76
post #40

Just tried K3 out for the first time today and it's a legitimate threat. Temporarily (maybe permanently) using it as my daily driver but it's wild how comparable it is to Opus 4.7/4.8 (what's been my go to for a bit now—wrote a quick post on what I found today [1]). [1] https://graybearding.bearblog.dev/kimi-k3-is-insane/

threat?

that's a weird way of describing a near frontier open weights un-crippled useful coding buddy

do you work for OpenAI or Anthropic per chance?

Re: Kimi K3 Architecture Overview and Notes

#77
post #67

Genuine question: how reproducible / usable / verifiable are these architectures from the published documentation? Are they similar to PDF/DWG/PSD specifications, where the format look like an open spec at first sight until you attempt to implement it and realize the crucial implementation details are undocumented?

It's entirely reproducible from the available documentation (which is why you see vLLM, SGLang, MLX etc all racing to produce optimized implementations). (As an aside, this is why the "open weights are not open source" thing is a complete misunderstanding. The weights themselves along with the documentation give you enough to fine tune the LLM. You can't rebuild it from scratch, but you can't do this even with the da…

Isn’t this a bit like saying “ ‘open object files are not open source’ is a complete misunderstanding” because “You can’t rebuild the executable from scratch, but you can’t do this even with the source code anyway (because of build nondeterminism / compiler versions / etc.)”?

Re: Kimi K3 Architecture Overview and Notes

#78

Genuine question: how reproducible / usable / verifiable are these architectures from the published documentation? Are they similar to PDF/DWG/PSD specifications, where the format look like an open spec at first sight until you attempt to implement it and realize the crucial implementation details are undocumented?

[flagged]

Re: Kimi K3 Architecture Overview and Notes

#79
post #60

Earlier quoted context omitted.

This is so true honestly.

Also Reddit-style humorous replies were frowned upon because they ruin serious discussion and encourage karma chasing as opposed to providing valuable insight.

This is why on Slashdot you could vote on insightful, humorous, etc. Something neither reddit, not hn have caugh on yet. Digg could have done it.
Post reply on HN