Either they have Blackwell with native 4-bit floating math, or they use have Chinese domestic NPU that support mxfp4 natively.
The article’s statement does not make sense.
11–20 of 79 posts
Either they have Blackwell with native 4-bit floating math, or they use have Chinese domestic NPU that support mxfp4 natively.
The article’s statement does not make sense.
Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture. In the recent leaked DeepSeek investor…
Genuine question. Is there a time factor part of that equation?
K3 is natively trained to mxfp4, if they cannot get a hold of Blackwell chip, it is meaningless. Hopper does not do native 4-bit floating math. Either they have Blackwell with native 4-bit floating math, or they use have Chinese domestic NPU that support mxfp4 natively. The article’s statement does not make sense.
It kind of confirms a hypothesis I have that the next phase of AI development will be about getting smaller (in terms of model size and compute), because smaller is more capital efficient for training (allowing faster iteration and more iteration cycles for a given amount of capital), allows for denser inference (more inference for a given amount of compute hardware), and allows for more edge inference applications.…
Getting smaller is how the last revolution in computing happened. A VAX 11/780 was good, but an 80386 was a lot better, since the latter could run on 3 AA batteries and the former needed 6,000 watts of 3 phase.
Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs). I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture. In the recent leaked DeepSeek investor…
It kind of confirms a hypothesis I have that the next phase of AI development will be about getting smaller (in terms of model size and compute), because smaller is more capital efficient for training (allowing faster iteration and more iteration cycles for a given amount of capital), allows for denser inference (more inference for a given amount of compute hardware), and allows for more edge inference applications.…
My thesis is still that model building has no moat. Folks continue to migrate around between the big labs. There is a lot of value in having good taste around the harness and how the models are used. The medium to long term winners will be the folks that control the compute.
Someone will find a way to make cheaper compute, and since nothing fundamental changed there (LLM didn't change how silicon were made), that's bound to happen.