Live data from Hacker News

Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

bloomberg.com

31–40 of 79 posts

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#31
post #7

It kind of confirms a hypothesis I have that the next phase of AI development will be about getting smaller (in terms of model size and compute), because smaller is more capital efficient for training (allowing faster iteration and more iteration cycles for a given amount of capital), allows for denser inference (more inference for a given amount of compute hardware), and allows for more edge inference applications.…

It begs an interesting question as well: are frontier labs stuck in the incumbent phase of the "innovator's dilemma," where there's such intense pressure to maintain the flavor/character of their widely-loved existing models, that they dare not invest their resources into radical reimagination of their approach towards architectures for smaller models? A company like Moonshot does not have this cultural limitation.

The frontier labs would be well served in carving out 20k sub-clusters and giving research teams carte blanche in building things with radically different architectures - with full permission to distill whatever they want from the flagship models. We'd expect to see more product lines that feel "different" from the flagship models if this were already being done.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#32

Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture. In the recent leaked DeepSeek investor…

everything about the current 'a.i' models - east vs west points the wastefulness of western labs - whether in terms of gpu clusters being ran, cost of inference.

the only way is down for the massive valuations and 'a.i' revenue projections.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#33
post #8

Earlier quoted context omitted.

Genuine question. Is there a time factor part of that equation?

It's obviously a question - did they just train for 10x as long due to having 10x fewer GPUs, but then that spoils the claim that they distilled Fable which was only recently introduced. Now doubt they did use some training data generated from older US models though.

This is just not how it works - it is perfectly plausible (in fact the most likely) that pretraining was well finished by the time Fable was released. A bit of extra distillation takes far less compute.

Moreover, it’s very plausible (and expected) to use multiple clusters and GPU types for RL rollouts which could very well not be included in this count.

No part of this pipeline is fixed in stone.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#34
post #11

K3 is natively trained to mxfp4, if they cannot get a hold of Blackwell chip, it is meaningless. Hopper does not do native 4-bit floating math. Either they have Blackwell with native 4-bit floating math, or they use have Chinese domestic NPU that support mxfp4 natively. The article’s statement does not make sense.

What do you mean by "it's meaningless" can you expand?

They mean that K3's native floating point is MXFP4, a newer standard that's not supported in Hopper, an older nvidia chipset, so it's impossible the GPUs in question were Hopper (the article specifically mentions H200s), which is correct.

Probably either the author or the people they interviewed got the specific GPU details mixed up or wrong. Maybe it was B200s.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#35
post #32

Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture. In the recent leaked DeepSeek investor…

everything about the current 'a.i' models - east vs west points the wastefulness of western labs - whether in terms of gpu clusters being ran, cost of inference. the only way is down for the massive valuations and 'a.i' revenue projections.

The wastefulness is akin to having multiple private jets on standby, or workers coming in to a fishbowl office to be gawked at, or massive datacenters to be used as social "proof".

Its all a fucking capitalistic farce to display to other rich elite that "Look at how much clout I have! I can make these peons dance around and do my bidding! Im a slave-owner!"

https://infosec.exchange/@david_chisnall/116991627711001827

He noted that that Silicon Valley doesnt really want to SOLVE problems. They want to find already-solved problems with problem matching. And of course, we just throw more people and more compute instead of optimization and understanding.

The Chinese are being actively constrained with bullshit politics around a second Red Scare moment. And, well, they're winning. A lot.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#36

Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture. In the recent leaked DeepSeek investor…

From the article:

> A person familiar with Moonshot’s procurement strategy confirmed that the company does indeed have a channel for accessing Blackwell processors via Southeast Asia. They didn’t specify whether this was a rental channel — which is legal, in most cases — or direct purchases, which are a breach of US regulations. The Information reported this week that Moonshot is seeking additional Blackwell processors to train its next model.

> The 20,000 chips Moonshot accesses via Alibaba, meanwhile, are from Nvidia’s earlier generation of Hopper products, the people familiar with the agreement said.

So the 20k GPUs from Alibaba is only a lower bound on how many you need to train a model like Kimi K3.

The advantage of having more GPUs in any case is not so much that you can train bigger models, but that the turnaround time is faster, so you can run more experiments to dial in training choices. It's entirely possible that Musk has more than enough compute, but can't hire the talent to run all those experiments. (That would also explain why he has excess capacity he can rent to Google.)

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#37
post #34

Earlier quoted context omitted.

What do you mean by "it's meaningless" can you expand?

They mean that K3's native floating point is MXFP4, a newer standard that's not supported in Hopper, an older nvidia chipset, so it's impossible the GPUs in question were Hopper (the article specifically mentions H200s), which is correct. Probably either the author or the people they interviewed got the specific GPU details mixed up or wrong. Maybe it was B200s.

Oh, okay, thank you for the explanation. I was under the impression that we, the United States, were banning exports of such GPUs.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#38
post #7

It kind of confirms a hypothesis I have that the next phase of AI development will be about getting smaller (in terms of model size and compute), because smaller is more capital efficient for training (allowing faster iteration and more iteration cycles for a given amount of capital), allows for denser inference (more inference for a given amount of compute hardware), and allows for more edge inference applications.…

Eventually, multi-modality might even get offloaded as workflows, which might allow models to do this at a fraction of the data and compute required.

I am adding multi-modality to https://github.com/guilt/TinyToT, and I see that dis-aggregating capabilities, very similar to how our own sensory organs work, seems to be paying off quite well.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#39
If Chinese companies are able to just reliably rent the damn GPUs then I really don't see how these restrictions are meant to be effective in the slightest. I guess it makes it less convenient and supply less certain, and retains some optionality for cutting off supply later?

If Alibaba can't bring the chips into China, but can buy a whole bunch in Thailand or Singapore or whereever (or secure exclusive rights via JV partners where relevant) and then just provide them as a service to its customers in China - what is the point? I'm sure many customers would actually prefer such an arrangement.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#40

Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture. In the recent leaked DeepSeek investor…

> they also mentioned only having a 20K GPU cluster (unclear if NVIDIA, or Huawei).

A few quotes from the transcript:

> Our current computing capacity is approximately 20,000 H-equivalent units, most of which have just arrived within the past month or two

> Regarding the Huawei 950, Huawei currently provides us with 16,000 SIM cards

> A Huawei 950 [cluster] with 16,000 cards is equivalent to only a B-series card [cluster] with 4,000 cards.

Post reply on HN