Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture. In the recent leaked DeepSeek investor…
Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba
41–50 of 79 posts
Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba
#42It's also built on the back of downloading all the content of the top ten sites on the list, the fact that all the independent content sites with their "pre-war" steel haven't been acquired at top dollar by U.S. companies shows the existential threat / stop China posturing is a total joke. https://archive.is/AvWWX
Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba
#43Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture. In the recent leaked DeepSeek investor…
(We've changed the title to what the article says now.)
Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba
#44It kind of confirms a hypothesis I have that the next phase of AI development will be about getting smaller (in terms of model size and compute), because smaller is more capital efficient for training (allowing faster iteration and more iteration cycles for a given amount of capital), allows for denser inference (more inference for a given amount of compute hardware), and allows for more edge inference applications.…
If they got a lot cheaper and only a little dumber, and we got a little smarter about how we use them, they could appear to the bystander as much more useful than they are.
Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba
#45If Chinese companies are able to just reliably rent the damn GPUs then I really don't see how these restrictions are meant to be effective in the slightest. I guess it makes it less convenient and supply less certain, and retains some optionality for cutting off supply later? If Alibaba can't bring the chips into China, but can buy a whole bunch in Thailand or Singapore or whereever (or secure exclusive rights via JV…
Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba
#46Earlier quoted context omitted.
It's obviously a question - did they just train for 10x as long due to having 10x fewer GPUs, but then that spoils the claim that they distilled Fable which was only recently introduced. Now doubt they did use some training data generated from older US models though.
This is just not how it works - it is perfectly plausible (in fact the most likely) that pretraining was well finished by the time Fable was released. A bit of extra distillation takes far less compute. Moreover, it’s very plausible (and expected) to use multiple clusters and GPU types for RL rollouts which could very well not be included in this count. No part of this pipeline is fixed in stone.
Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba
#47Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture. In the recent leaked DeepSeek investor…
This feels like a very American way of designing things - just throw more horse power at it, bigger is better! The rest of the world is usually a bit more resource constrained and efficient at using those resources.
See also Mustangs vs German sports cars, giant American fridges, giant American suburban McMansions vs livable cities etc etc.
Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba
#48I read this as if they purchased it from Alibaba the marketplace.....
Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba
#49https://archive.is/vgmMB
Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba
#50Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture. In the recent leaked DeepSeek investor…
It's not that surprising to me. Most of the innovation in Chinese models has been in efficiency gains and optimizations. K3 coming from the factory in MXFP4 weights is a pretty relevant factor. Big performance gap probably also due to Moonshot doing QAT. Throw in the fact that Musk has easier access to compute, and I think you have your answer on the disparity.
In this sense, any advance in intelligence is a performance improvement and vice versa