Live data from Hacker News

Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

bloomberg.com

51–60 of 79 posts

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#51
post #32

Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture. In the recent leaked DeepSeek investor…

everything about the current 'a.i' models - east vs west points the wastefulness of western labs - whether in terms of gpu clusters being ran, cost of inference. the only way is down for the massive valuations and 'a.i' revenue projections.

> west points the wastefulness of western labs whether in terms of gpu clusters being ran, cost of inference.

Can’t possibly be intentional media strategy by a geopolitical target, right? If its going to negatively impact valuations and revenue of major rivals, seems like a desirable strategy?

We just saw OpenAI take the steps to significantly lower the cost of one of their models, which confirms that at least one western lab has a large margin on inference, not a large inference cost.

Meanwhile, we don’t know how many experiments these east/west labs are performing relative to each other. We also know that many western labs have a whole portfolio of models too, which is product breadth not necessarily waste.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#52
post #35
post #32

Earlier quoted context omitted.

everything about the current 'a.i' models - east vs west points the wastefulness of western labs - whether in terms of gpu clusters being ran, cost of inference. the only way is down for the massive valuations and 'a.i' revenue projections.

The wastefulness is akin to having multiple private jets on standby, or workers coming in to a fishbowl office to be gawked at, or massive datacenters to be used as social "proof". Its all a fucking capitalistic farce to display to other rich elite that "Look at how much clout I have! I can make these peons dance around and do my bidding! Im a slave-owner!" https://infosec.exchange/@david_chisnall/116991627711001827…

No, it's an attempt to form an artificial moat by buying all the world's GPUs and memory.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#53

Earlier quoted context omitted.

It's obviously a question - did they just train for 10x as long due to having 10x fewer GPUs, but then that spoils the claim that they distilled Fable which was only recently introduced. Now doubt they did use some training data generated from older US models though.

This is just not how it works - it is perfectly plausible (in fact the most likely) that pretraining was well finished by the time Fable was released. A bit of extra distillation takes far less compute. Moreover, it’s very plausible (and expected) to use multiple clusters and GPU types for RL rollouts which could very well not be included in this count. No part of this pipeline is fixed in stone.

OK, but that doesn't address the question of what can be inferred from Moonshot apparently having access to far less compute then the western labs. To what extent does this reflect the need for less compute due to the (published) architectural efficiencies of their model, and to what extent that they just trained for longer (esp. pre-training - 80% of total compute, perhaps)?

I think the word "distillation" needs to be used a bit more selectively here. If their pre-training run was complete before Fable was released that implies that ZERO Fable data went into the base model. Perhaps the timeline allows for a few weeks at best of incremental post-training on some limited amount of Fable data, but calling this "distillation" seems a bit dramatic especially given the redacted outputs that would have been available. A more factual speculation would just be that they may have had time to post-train using a limited amount of Fable output in some fashion (LLM as judge? SFT? Who knows ...).

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#54
post #47

Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture. In the recent leaked DeepSeek investor…

> Struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). This feels like a very American way of designing things - just throw more horse power at it, bigger is better! The rest of the world is usually a bit more resource constrained and efficient at using those resources. See also Mustangs vs German sports cars, giant American fridges, giant…

You've just got fridge jealousy! ;-)

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#55

Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture. In the recent leaked DeepSeek investor…

It's not that surprising to me. Most of the innovation in Chinese models has been in efficiency gains and optimizations. K3 coming from the factory in MXFP4 weights is a pretty relevant factor. Big performance gap probably also due to Moonshot doing QAT. Throw in the fact that Musk has easier access to compute, and I think you have your answer on the disparity.

I don't follow this. Clearly all of the frontier labs are doing these things.

When OAI released gpt-oss it was released as an mxfp4 checkpoint.

OAI, Ant, et al are also obviously employing QAT.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#56
post #35

Earlier quoted context omitted.

The wastefulness is akin to having multiple private jets on standby, or workers coming in to a fishbowl office to be gawked at, or massive datacenters to be used as social "proof". Its all a fucking capitalistic farce to display to other rich elite that "Look at how much clout I have! I can make these peons dance around and do my bidding! Im a slave-owner!" https://infosec.exchange/@david_chisnall/116991627711001827…

No, it's an attempt to form an artificial moat by buying all the world's GPUs and memory.

True. But check out Ebay right now. Networking gear is CHEAP.

I just bought 2 switches, 48 port 10GbE with 4x QSFP+ at 40Gb fiber. $110 each.

You can even get 24 port QSFP+ @100Gb networking devices for $350.

Yeah while ram and gfx is $$$$$, networking is rock bottom prices..

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#57
post #7

It kind of confirms a hypothesis I have that the next phase of AI development will be about getting smaller (in terms of model size and compute), because smaller is more capital efficient for training (allowing faster iteration and more iteration cycles for a given amount of capital), allows for denser inference (more inference for a given amount of compute hardware), and allows for more edge inference applications.…

> It kind of confirms a hypothesis I have that the next phase of AI development will be about getting smaller

isn't that, in the end, the case with all/most technologies?

i can get a petaflop of compute capacity in a dgx spark for relatively cheap nowadays. that used to be a whole supercomputer like 20 years ago.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#58
post #34

Earlier quoted context omitted.

What do you mean by "it's meaningless" can you expand?

They mean that K3's native floating point is MXFP4, a newer standard that's not supported in Hopper, an older nvidia chipset, so it's impossible the GPUs in question were Hopper (the article specifically mentions H200s), which is correct. Probably either the author or the people they interviewed got the specific GPU details mixed up or wrong. Maybe it was B200s.

Notably, MXFP4 was introduced at the (much less costly) supervised fine-tuning stage after pretraining, so the number of B200/B300 GPUs could be relatively small in comparison to the number of H200 GPUs used during pretraining (or maybe not, who knows).

(Kimi K3 tech report section 4.1.1 https://arxiv.org/pdf/2607.24653)

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#59
post #21

Earlier quoted context omitted.

Sounds fair, since Sam can run a for profit NGO

I’m still baffled that OpenAI can complain about this with a straight face while these models are literally trained on everything regardless of copyright or permission.

Raw stolen data that is in no way related to AI vs a very very expensive transformative compilation of that raw stolen data that results in usable AI. The frustration is completely understandable to me, since distillation skips much of that "very very expensive" part.

And yes, I understand the stolen data was expensive to make, so I understand the owners of it are also frustrated, but that's partly a problem with current law. Would the authors of the world be rich if OpenAI bought a single copy of their book to legally scan? For best sellers, that's somewhere around pennies, so no. Should the authors get a share in OpenAI? Current laws says, unambiguously, "no".

Frustration all around is reasonable.

Re: Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba

#60
post #39

If Chinese companies are able to just reliably rent the damn GPUs then I really don't see how these restrictions are meant to be effective in the slightest. I guess it makes it less convenient and supply less certain, and retains some optionality for cutting off supply later? If Alibaba can't bring the chips into China, but can buy a whole bunch in Thailand or Singapore or whereever (or secure exclusive rights via JV…

"what is the point?" There is no point, just zero-sum thinking which isn't compatible with reality.

Zero-sum thinking that everyone is thinking [1].

[1] https://www.tomshardware.com/tech-industry/artificial-intell...

Post reply on HN