License: https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE > If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or…
Kimi 2.6 had the same janky license: https://huggingface.co/moonshotai/Kimi-K2.6/blob/main/LICENS... - looks like they've been doing that at least as far back as K2. The models they released in 2025 - https://huggingface.co/moonshotai/models - were clean MIT. They started doing the "modified MIT" thing in January 2026 with moonshotai/Kimi-K2-Thinking
Kimi-K3 Technical Report [pdf]
61–70 of 199 posts
Re: Kimi-K3 Technical Report [pdf]
#62It's funny that we've finally returned to tanh activation functions, time is a circle.
Re: Kimi-K3 Technical Report [pdf]
#63Re: Kimi-K3 Technical Report [pdf]
#64Earlier quoted context omitted.
This comment would be much better without the second line
I'm not sure I understand the case for open-source models being decelerationist, is this it? Decel: - Potentially reduces investor appetite for funding big labs. - More risk of powerful AI getting in bad hands -> more regulation. Accel: - More competition so big labs can't rest on laurels. - More research in open, so all labs can accrete advancements faster. I feel like open-source = acceleration has a much more clea…
Re: Kimi-K3 Technical Report [pdf]
#65Also open sourced a bunch of infra to go with it. Anyone who claims open source and open weights models are "decel" needs to get their head checked https://github.com/MoonshotAI/MoonEP https://github.com/kvcache-ai/AgentEnv https://github.com/MoonshotAI/FlashKDA
This comment would be much better without the second line
Re: Kimi-K3 Technical Report [pdf]
#66What would it take to get an American open model to compete with this?
Latest "big" release from any of the bigger American lab must have been GPT-OSS-120b I think? Released ~summer 2025, so pretty much one years ago. Doesn't seem like it'll happen by itself, so something either forcing their hand figuratively, or something forcing their hand literally. Personally I was wishing/hoping for one of the recent Gemma releases to be in the ~100B class at least, but sadly Google is keeping tha…
Re: Kimi-K3 Technical Report [pdf]
#67Also open sourced a bunch of infra to go with it. Anyone who claims open source and open weights models are "decel" needs to get their head checked https://github.com/MoonshotAI/MoonEP https://github.com/kvcache-ai/AgentEnv https://github.com/MoonshotAI/FlashKDA
the only reason other labs can catch up is because the frontier labs can be distilled, and they siphon a % of the labs' revenue to reinvest into the next iteration
full accel would mean nationalizing the big 2 labs and locking in manhattan project style until RSI
(Edit: some great counterpoints in the replies. my view has definitely been changed!)
Re: Kimi-K3 Technical Report [pdf]
#68Re: Kimi-K3 Technical Report [pdf]
#69Back of the envelope calculation (could be off, correct me if I am) If you are a large enough company that spends million+ on inference a month, it makes sense to buy a GB300 rack ($6M on top range from what I could find) which has 20.7 TB. Since the model is mixed trained (MXFP4), you would need less than 10% of the rack's memory to serve the full model. Aggregate HBM bandwidth: 576 TB/s. You can run over 6000 paral…
Re: Kimi-K3 Technical Report [pdf]
#70Back of the envelope calculation (could be off, correct me if I am) If you are a large enough company that spends million+ on inference a month, it makes sense to buy a GB300 rack ($6M on top range from what I could find) which has 20.7 TB. Since the model is mixed trained (MXFP4), you would need less than 10% of the rack's memory to serve the full model. Aggregate HBM bandwidth: 576 TB/s. You can run over 6000 paral…
That another 400 to 700k.
It becomes your problem and not someone else’s. However, I don’t trust hosted LLMs for anything that needs to be private.