Live data from Hacker News

Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

bloomberg.com

41–50 of 79 posts

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#41

Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture. In the recent leaked DeepSeek investor…

Do we believe them? It seems there’s no reason with the resources OpenAI and Anthropic have they wouldn’t be incentivised to build equally efficient pipelines? I think it’s mostly the fine tuning with distillation that gives Chinese labs their big advantage. However I’m pretty sure all the main labs are poisoning the output now when they detect Chinese activity so while K3 looks frontier on tests when you use it it’s miles off Sol and Fable.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#42

It's also built on the back of downloading all the content of the top ten sites on the list, the fact that all the independent content sites with their "pre-war" steel haven't been acquired at top dollar by U.S. companies shows the existential threat / stop China posturing is a total joke. https://archive.is/AvWWX

No one is using that dataset for any serious models. It’s way too small and old.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#43

Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture. In the recent leaked DeepSeek investor…

From a superficial look I think the submitted title ("Moonshot built on 20k Nvidia chip cluster from Alibaba") might have been misleading. Article says "That hardware forms a substantial chunk of the overall computing capacity Moonshot uses for its Kimi models", implying that other hardware was also used.

(We've changed the title to what the article says now.)

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#44
post #7

It kind of confirms a hypothesis I have that the next phase of AI development will be about getting smaller (in terms of model size and compute), because smaller is more capital efficient for training (allowing faster iteration and more iteration cycles for a given amount of capital), allows for denser inference (more inference for a given amount of compute hardware), and allows for more edge inference applications.…

It feels like we're not really bottlenecked by model performance now anyway. Most cases I find where a model is being an idiot can be remedied by rethinking the context and tooling available to that model.

If they got a lot cheaper and only a little dumber, and we got a little smarter about how we use them, they could appear to the bystander as much more useful than they are.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#45
post #39

If Chinese companies are able to just reliably rent the damn GPUs then I really don't see how these restrictions are meant to be effective in the slightest. I guess it makes it less convenient and supply less certain, and retains some optionality for cutting off supply later? If Alibaba can't bring the chips into China, but can buy a whole bunch in Thailand or Singapore or whereever (or secure exclusive rights via JV…

"what is the point?" There is no point, just zero-sum thinking which isn't compatible with reality.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#46

Earlier quoted context omitted.

It's obviously a question - did they just train for 10x as long due to having 10x fewer GPUs, but then that spoils the claim that they distilled Fable which was only recently introduced. Now doubt they did use some training data generated from older US models though.

This is just not how it works - it is perfectly plausible (in fact the most likely) that pretraining was well finished by the time Fable was released. A bit of extra distillation takes far less compute. Moreover, it’s very plausible (and expected) to use multiple clusters and GPU types for RL rollouts which could very well not be included in this count. No part of this pipeline is fixed in stone.

Relatedly I also wonder how much the orchestrated distillation bot nets that Anthropic uses against China are real. Not saying it does not happen but I have never seen dating showing how pervasive it is and I always imagine other western labs would absolutely be doing the same.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#47

Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture. In the recent leaked DeepSeek investor…

> Struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs).

This feels like a very American way of designing things - just throw more horse power at it, bigger is better! The rest of the world is usually a bit more resource constrained and efficient at using those resources.

See also Mustangs vs German sports cars, giant American fridges, giant American suburban McMansions vs livable cities etc etc.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#50

Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture. In the recent leaked DeepSeek investor…

It's not that surprising to me. Most of the innovation in Chinese models has been in efficiency gains and optimizations. K3 coming from the factory in MXFP4 weights is a pretty relevant factor. Big performance gap probably also due to Moonshot doing QAT. Throw in the fact that Musk has easier access to compute, and I think you have your answer on the disparity.

But perf improvements not only mean you can run the same thing on cheaper hardware, you can also run more capable models in the same hardware

In this sense, any advance in intelligence is a performance improvement and vice versa

Post reply on HN