Live data from Hacker News

Running Kimi K3 on MI355X at Better Performance per Dollar Than B300

wafer.ai

91–100 of 119 posts

Re: Running Kimi K3 on MI355X at Better Performance per Dollar Than B300

#93

This part sounds like AI assisted setting this up and benchmarking it: >The fix was trivially simple: zero-pad the head count 12→16, run the fast kernel, and extract the real 12 heads from the output. I've recently used a frontier AI (ChatGPT 5.6 Sol on ultra) to set up a much smaller local model, and the performance optimizations it introduced left the model totally incoherent. (The model just repeats a single chara…

hi I work at wafer. yes we ran benchmarks. for example our Kimi K3 is live on open router and in order to host there you have to run accuracy checks. like tau/gpqa.

and then u must pass test regarding thinking, coherency, and tool calling

Re: Running Kimi K3 on MI355X at Better Performance per Dollar Than B300

#94
Another ad posing as "science".

It's basically a Wafer/AMD advertisement.

B300 is about 46% faster for one stream and 65% faster in aggregate. AMD wins only after Wafer divides throughput // selected cloud-rental prices: $2.50/GPU-hour for MI355X versus $6 for B300.

Benchmarks are unreproducible, power costs are missing, ROCm was patched... and on and on.

Sure, it can serve that particular model at that particular size more economically. Good for them... in particular.

Re: Running Kimi K3 on MI355X at Better Performance per Dollar Than B300

#95

This part sounds like AI assisted setting this up and benchmarking it: >The fix was trivially simple: zero-pad the head count 12→16, run the fast kernel, and extract the real 12 heads from the output. I've recently used a frontier AI (ChatGPT 5.6 Sol on ultra) to set up a much smaller local model, and the performance optimizations it introduced left the model totally incoherent. (The model just repeats a single chara…

I have never used sol, but I regularly use Claude to set up and benchmark local models per task. It is very thorough and has always returned good setups. The only thing you have to do is make sure you point at the model card. It’s always incredulous that models exist after its training cutoff. K3 does an ok job of setting up, but its config searching isn’t nearly as thorough and its will confidently tell you it’s fou…

>and its will confidently tell you it’s found the best setup when it’s only turned a few knobs.

on one hand , yeah : a model being more aggressive towards exploration of the decision space is usually a good thing.

on the other hand : I think that it's up to the operator to set rigid test criteria to make these things actually work well in a repeatable fashion.

so in other words, i'm glad claude is doing a better job for the way you're prompting the thing, but as a spectator from afar these kind of operator complaints usually spring up from the use of weak, ambiguous, under-considered prompts.

similar with the parents' complaint; there should have been a testing criteria for legible output that got immediately flagged or failed by the larger model.

it's a very hard sell for me to think that " a a a a a a a " is accepted as a valid language output by any near-SOTA-large model without some real coercion.

Re: Running Kimi K3 on MI355X at Better Performance per Dollar Than B300

#97

Another ad posing as "science". It's basically a Wafer/AMD advertisement. B300 is about 46% faster for one stream and 65% faster in aggregate. AMD wins only after Wafer divides throughput // selected cloud-rental prices: $2.50/GPU-hour for MI355X versus $6 for B300. Benchmarks are unreproducible, power costs are missing, ROCm was patched... and on and on. Sure, it can serve that particular model at that particular si…

Since we're (as a society) spending tens of billions of dollars to serve models like this, it's a pretty interesting advertisement.

Re: Running Kimi K3 on MI355X at Better Performance per Dollar Than B300

#98

Earlier quoted context omitted.

thanks for sharing. When I read your original comment, I was thinking you had just asked it to evaluate the article. (Like just "evaluate this article" or something.) I don't think anything anyone (or any AI) has ever written or published (including Sol itself) wouldn't be torn apart by the prompt you gave though.

I think it's fair to expect an extensive review of an article before publishing. Not everything has a set of serious flaws.

>Not everything has a set of serious flaws.

I'm saying everything looks like it does, with the prompt you gave Sol. Go ahead and point the same prompt to anything you don't think has serious flaws and you'll see it would tear it apart.

Re: Running Kimi K3 on MI355X at Better Performance per Dollar Than B300

#100
post #95

Earlier quoted context omitted.

I have never used sol, but I regularly use Claude to set up and benchmark local models per task. It is very thorough and has always returned good setups. The only thing you have to do is make sure you point at the model card. It’s always incredulous that models exist after its training cutoff. K3 does an ok job of setting up, but its config searching isn’t nearly as thorough and its will confidently tell you it’s fou…

>and its will confidently tell you it’s found the best setup when it’s only turned a few knobs. on one hand , yeah : a model being more aggressive towards exploration of the decision space is usually a good thing. on the other hand : I think that it's up to the operator to set rigid test criteria to make these things actually work well in a repeatable fashion. so in other words, i'm glad claude is doing a better job…

It’s not a complaint, only an observation. Operator aside, to observe that the models behave differently shouldn’t be controversial, should it?

It seems like you read my comment as dissing Kimi. Kimi is in regular rotation for me. Some tasks that Claude used to own go to Kimi first now. They are just different tools with their own strengths and weaknesses.

Post reply on HN