Live data from Hacker News

Running Kimi K3 on MI355X at Better Performance per Dollar Than B300

wafer.ai

111–119 of 119 posts

Re: Running Kimi K3 on MI355X at Better Performance per Dollar Than B300

#111

Earlier quoted context omitted.

> The Chinese models are mostly very well documented in terms of architecture and training processes/flows, with what is missing to recreate them being the training data. ...and because that training data is missing, they can't be replicated. Which means that you cannot assert that the Chinese are being open in their LLM development, because there's no way to verify that the techniques they describe are actually the…

You can replicate the architectural innovations, and try them for yourself with your own dataset. It seems some of them are certainly being used by western companies, such as DeepSeek Sparse Attention, now supported by NVIDIA cuDNN. Ditto for training algorithms and procedures such as Slime or DeepSeek's details instructions on how to build a reasoning model. This is the exact value of openly shared details - others…

> You can replicate the architectural innovations, and try them for yourself with your own dataset.

That's not related to my comment. My comment was pointing out that you can't verify something that wasn't published. You have no idea what fraction of their techniques they're not publishing, and how much they contribute to their model performance, because you cannot replicate the models, because they don't publish their training data.

> and the reason the Chinese are not sharing data are no more nefarious than why the American companies are not sharing

This is moving the goalposts. Your claim was that "The Chinese have actually been very open about training", which is false, as discussed. Nobody ever claimed that the American labs were open.

Re: Running Kimi K3 on MI355X at Better Performance per Dollar Than B300

#113

Earlier quoted context omitted.

You can replicate the architectural innovations, and try them for yourself with your own dataset. It seems some of them are certainly being used by western companies, such as DeepSeek Sparse Attention, now supported by NVIDIA cuDNN. Ditto for training algorithms and procedures such as Slime or DeepSeek's details instructions on how to build a reasoning model. This is the exact value of openly shared details - others…

> You can replicate the architectural innovations, and try them for yourself with your own dataset. That's not related to my comment. My comment was pointing out that you can't verify something that wasn't published. You have no idea what fraction of their techniques they're not publishing, and how much they contribute to their model performance, because you cannot replicate the models, because they don't publish the…

> You have no idea what fraction of their techniques they're not publishing, and how much they contribute to their model performance, because you cannot replicate the models, because they don't publish their training data.

Why would you be concerned about THEIR model performance ?!

Surely if you are an ML researcher and read about a new technique, you are interested in how YOU may be able to use it.

If you are Ilya Sutskever sitting at OpenAI in 2017, and happen upon Google's "attention" (Transformer architecture) paper, then what you do is go and implement it for yourself, get yourself some training data, and try it.

What you are NOT going to do is whine about not being given their source code, or their training data, or their training harness, or a dump of Google's corporate secrets. You take the research that has been shared and evaluate it for yourself.

Re: Running Kimi K3 on MI355X at Better Performance per Dollar Than B300

#114

Earlier quoted context omitted.

> You can replicate the architectural innovations, and try them for yourself with your own dataset. That's not related to my comment. My comment was pointing out that you can't verify something that wasn't published. You have no idea what fraction of their techniques they're not publishing, and how much they contribute to their model performance, because you cannot replicate the models, because they don't publish the…

> You have no idea what fraction of their techniques they're not publishing, and how much they contribute to their model performance, because you cannot replicate the models, because they don't publish their training data. Why would you be concerned about THEIR model performance ?! Surely if you are an ML researcher and read about a new technique, you are interested in how YOU may be able to use it. If you are Ilya S…

This doesn't have anything to do with Google or OpenAI. I'm not Ilya Sutskever and it's not 2017, either.

I'm not "whining" about anything. You made the claim "The Chinese have actually been very open about training" and I showed that that was false. That's all that there is to it.

Re: Running Kimi K3 on MI355X at Better Performance per Dollar Than B300

#115

Earlier quoted context omitted.

> You have no idea what fraction of their techniques they're not publishing, and how much they contribute to their model performance, because you cannot replicate the models, because they don't publish their training data. Why would you be concerned about THEIR model performance ?! Surely if you are an ML researcher and read about a new technique, you are interested in how YOU may be able to use it. If you are Ilya S…

This doesn't have anything to do with Google or OpenAI. I'm not Ilya Sutskever and it's not 2017, either. I'm not "whining" about anything. You made the claim "The Chinese have actually been very open about training" and I showed that that was false. That's all that there is to it.

The rest of the word is reading Chinese published research and benefiting from it.

Apparently you are unaware of it and not benefiting from it.

Oh well.

Re: Running Kimi K3 on MI355X at Better Performance per Dollar Than B300

#116
Another proof point for AMD inference economics: Wafer reports running Kimi K3 on AMD MI355X GPUs with better performance per dollar than NVIDIA B300.

The important detail is memory. Kimi K3 requires more than 1.5TB of VRAM. With 288GB per GPU the MI355X can fit the model within one 8 GPU node.

This is exactly why Sciforium built its production inference stack on AMD.

We own & operate our AMD GPUs US data centers & serving layer. Our stack is already optimized so customers get a simple API rather than an AMD engineering project.

By removing the cloud GPU rental layer & provider markup we gain greater control over capacity cost & performance.

Try Sciforium On Demand & receive 25% off plus $200 in promotional credits with code AAI2026, sign up at https://console.sciforium.com/

Re: Running Kimi K3 on MI355X at Better Performance per Dollar Than B300

#117
Another proof point for AMD inference economics: Wafer reports running Kimi K3 on AMD MI355X GPUs with better performance per dollar than NVIDIA B300.

The key detail is memory. Kimi K3 requires more than 1.5TB of VRAM. With 288GB per GPU the MI355X can fit the model within one 8 GPU node.

This is exactly why Sciforium built its production inference stack on AMD. We own & operate our AMD GPUs US data centers & serving layer. Our stack is already optimized so customers get a simple API rather than an AMD engineering project.

By removing the cloud GPU rental layer & provider markup we gain greater control over capacity cost & performance.

Try Sciforium On Demand & receive 25% off plus $200 in promotional credits with code AAI2026, sign up at https://console.sciforium.com/

Re: Running Kimi K3 on MI355X at Better Performance per Dollar Than B300

#118
post #37
post #35

Earlier quoted context omitted.

If you would tell people at the start of 20th century how much energy we consume, they may not believe you or think it is wasteful.

Looks like global energy consumption has risen by an order of magnitude from 1900 to 2000: https://www.encyclopedie-energie.org/en/world-energy-consump... Unfortunately, electricity prices did not fall by the same factor, so I fear that training a frontier model will still cause a an unsustainable dent in my monthly budget.

Only an order of magnitude? I'm both surprised by both the fact it's so low and the fact 1900 is 3x of 1800.

Re: Running Kimi K3 on MI355X at Better Performance per Dollar Than B300

#119

Earlier quoted context omitted.

> Not to mention no one serious is serving this on 8xB200 instead of multiple nodes: the vast majority of Moonshot's inference work is focused on PD-disaggregation The GPU price discourse is absurd, but many are serving models on single node setups when the model fits

You can't beat current API pricing running Kimi K3 a single node, so as I said anyone serious is not doing that for Kimi K3 . Not sure how you're getting "no one serves models on a single node" out of that.

I meant serious players often do serve on a single node. They can beat API pricing as well. Multi-node can add gain, but also adds a lot of deployment complexity so its just not always possible or optimal.
Post reply on HN