Lol, such a lazily written article by wafer.ai GPUs. 8× MI355X (TP8) B300 (TP8+DCP8) Decode tok/s per stream 118 tok/s 172 tok/s Peak aggregate. 952 tok/s 1,568 tok/s Peak aggregate per GPU 119 tok/s 196 tok/s On every row the B300 beat the MI355X The B200 is being forcefully compared against something which is not gonna fit within it's memory in a single node & not much details about multi-node interconnectivity, di…
To the B200’s defence, its numbers are somewhat deflated by the fact that it pays a cross-node all-reduce on the decode critical path (RoCE v2 at ~195 Gb/s) — it’s the only config here that spans two nodes