Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba
1–10 of 79 posts
Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba
#2I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture.
In the recent leaked DeepSeek investor meeting, they also mentioned only having a 20K GPU cluster (unclear if NVIDIA, or Huawei).
Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba
#3I wonder how much spending is motivated by the phenomenon of sudden emergent performance in LLMs. Clearly some people who are smarter than me expect something like emergent AGI, or at least they think the odds justify spending whatever it takes to see if that would happen.
That leaves a lot of room for efficient aggressive followers.
Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba
#4Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture. In the recent leaked DeepSeek investor…
Source?
Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba
#5Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture. In the recent leaked DeepSeek investor…
> Musk (who freely admits to distilling OpenAI's models) Source?
Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba
#6Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba
#7The goal will be to develop smaller models with more efficient architectures, that have similar or even better performance than larger models.
Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba
#8Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture. In the recent leaked DeepSeek investor…
Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba
#9Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba
#10It kind of confirms a hypothesis I have that the next phase of AI development will be about getting smaller (in terms of model size and compute), because smaller is more capital efficient for training (allowing faster iteration and more iteration cycles for a given amount of capital), allows for denser inference (more inference for a given amount of compute hardware), and allows for more edge inference applications.…
A VAX 11/780 was good, but an 80386 was a lot better, since the latter could run on 3 AA batteries and the former needed 6,000 watts of 3 phase.