Live data from Hacker News

Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

bloomberg.com

1–10 of 79 posts

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#2
Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs).

I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture.

In the recent leaked DeepSeek investor meeting, they also mentioned only having a 20K GPU cluster (unclear if NVIDIA, or Huawei).

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#3
I would not minimize the extent to which AI progress depends on things like the effort that goes into training material acquisition and data labeling. In those areas, improvement is a direct function of how much you invest.

I wonder how much spending is motivated by the phenomenon of sudden emergent performance in LLMs. Clearly some people who are smarter than me expect something like emergent AGI, or at least they think the odds justify spending whatever it takes to see if that would happen.

That leaves a lot of room for efficient aggressive followers.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#4

Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture. In the recent leaked DeepSeek investor…

> Musk (who freely admits to distilling OpenAI's models)

Source?

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#5
post #4

Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture. In the recent leaked DeepSeek investor…

> Musk (who freely admits to distilling OpenAI's models) Source?

https://www.forbes.com/sites/antoniopequenoiv/2026/04/30/elo...

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#7
It kind of confirms a hypothesis I have that the next phase of AI development will be about getting smaller (in terms of model size and compute), because smaller is more capital efficient for training (allowing faster iteration and more iteration cycles for a given amount of capital), allows for denser inference (more inference for a given amount of compute hardware), and allows for more edge inference applications.

The goal will be to develop smaller models with more efficient architectures, that have similar or even better performance than larger models.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#8

Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture. In the recent leaked DeepSeek investor…

Genuine question. Is there a time factor part of that equation?

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#9
My thesis is still that model building has no moat. Folks continue to migrate around between the big labs. There is a lot of value in having good taste around the harness and how the models are used. The medium to long term winners will be the folks that control the compute.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#10
post #7

It kind of confirms a hypothesis I have that the next phase of AI development will be about getting smaller (in terms of model size and compute), because smaller is more capital efficient for training (allowing faster iteration and more iteration cycles for a given amount of capital), allows for denser inference (more inference for a given amount of compute hardware), and allows for more edge inference applications.…

Getting smaller is how the last revolution in computing happened.

A VAX 11/780 was good, but an 80386 was a lot better, since the latter could run on 3 AA batteries and the former needed 6,000 watts of 3 phase.

Post reply on HN