Earlier quoted context omitted.
This makes no sense, the V100 has more memory bandwidth than both the TPU and TPUv2
V100 has 900gb/s memory bandwidth [0]. TPUv2 has 600gb/s per chip x 4 chips, so 2400gb/s [1]. As we've discussed elsewhere [2], comparing TPUv2 to V100 on a per chip basis doesn't make much sense. Who cares how many chips are on the board? If Google announced tomorrow that TPUv3 is coming out, which is identical to TPUv2 but the four chips are glued together, nobody would care. The questions that we should instead be…
Nobody is comparing DGX1-V to a single TPUv2 chip, because it doesn't make any sense to do so. they are totally different kinds of machines. But for some reason everyone is comparing a cluster of 4 TPUv2 chips to a single V100 chip.
It only makes sense to compare 4xTPUv2 to 1xV100 if they are equivalent in some meaningful metric, like total die size, power, etc.
In lieu of any available data, I'm going to continue to assume that each TPUv2 chip is roughly comparable in terms of power & die size to each V100 chip. If this was grossly wrong, I would expect that all four would be condensed into a single chip, which would dramatically increase the performance of the interconnects.
We could resolve this rapidly if there were any data available about die size, TDP, anything of TPUv2.