Live data from Hacker News

Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

bloomberg.com

11–20 of 79 posts

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#11
K3 is natively trained to mxfp4, if they cannot get a hold of Blackwell chip, it is meaningless. Hopper does not do native 4-bit floating math.

Either they have Blackwell with native 4-bit floating math, or they use have Chinese domestic NPU that support mxfp4 natively.

The article’s statement does not make sense.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#12
post #8

Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture. In the recent leaked DeepSeek investor…

Genuine question. Is there a time factor part of that equation?

It's obviously a question - did they just train for 10x as long due to having 10x fewer GPUs, but then that spoils the claim that they distilled Fable which was only recently introduced. Now doubt they did use some training data generated from older US models though.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#13
post #4

Earlier quoted context omitted.

> Musk (who freely admits to distilling OpenAI's models) Source?

https://www.forbes.com/sites/antoniopequenoiv/2026/04/30/elo...

Sounds fair, since Sam can run a for profit NGO

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#16
post #11

K3 is natively trained to mxfp4, if they cannot get a hold of Blackwell chip, it is meaningless. Hopper does not do native 4-bit floating math. Either they have Blackwell with native 4-bit floating math, or they use have Chinese domestic NPU that support mxfp4 natively. The article’s statement does not make sense.

What do you mean by "it's meaningless" can you expand?

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#17
post #7

It kind of confirms a hypothesis I have that the next phase of AI development will be about getting smaller (in terms of model size and compute), because smaller is more capital efficient for training (allowing faster iteration and more iteration cycles for a given amount of capital), allows for denser inference (more inference for a given amount of compute hardware), and allows for more edge inference applications.…

Getting smaller is how the last revolution in computing happened. A VAX 11/780 was good, but an 80386 was a lot better, since the latter could run on 3 AA batteries and the former needed 6,000 watts of 3 phase.

Yes but that's because of the scaling laws for transistors. ML models seem to get better the bigger they are. If you want to compress the world's information, you need to look at all the information in the world.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#18

Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture. In the recent leaked DeepSeek investor…

It's not that surprising to me. Most of the innovation in Chinese models has been in efficiency gains and optimizations. K3 coming from the factory in MXFP4 weights is a pretty relevant factor. Big performance gap probably also due to Moonshot doing QAT. Throw in the fact that Musk has easier access to compute, and I think you have your answer on the disparity.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#19
post #7

It kind of confirms a hypothesis I have that the next phase of AI development will be about getting smaller (in terms of model size and compute), because smaller is more capital efficient for training (allowing faster iteration and more iteration cycles for a given amount of capital), allows for denser inference (more inference for a given amount of compute hardware), and allows for more edge inference applications.…

This is likely what Google has been doing, as it suites them best to have small fast models rather than hulking slow giants. Lower intelligence but way more ability to serve. Anthropic would probably need a datacenter the size of a small country to serve Fable on a Google search/services scale.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#20
post #9

My thesis is still that model building has no moat. Folks continue to migrate around between the big labs. There is a lot of value in having good taste around the harness and how the models are used. The medium to long term winners will be the folks that control the compute.

You can only control the compute if it is not commoditized, which at the current moment, it is not, leading to industry players having 75-80% margins.

Someone will find a way to make cheaper compute, and since nothing fundamental changed there (LLM didn't change how silicon were made), that's bound to happen.

Post reply on HN