Earlier quoted context omitted.
Its incomparable though. The Pegasus has far more compute power. Just because the Pegasus has 500W worst-case TDP doesn't mean that its average case would be 500W constant. If you scale back your code and idle parts of the GPU, you can drop the energy cost arbitrarily. At least, that's how GPUs on desktops work. They only use a ton of power if you give them a ton of work. Write your code in an energy-efficient manner…
Pegasus's chip is more general purpose and doubtless has far more general purpose compute power. But that's irrelevant. What's relevant is Tesla's chip is optimized specifically for their NN pipeline whereas Pegasus is based off of general purpose GPU architecture, thus Tesla's chip achieves a better TOPS per Watt than Pegasus. And it'd be strange if it didn't. "Tesla's chip will NEVER be able to scale above 100W" ok…
Theoretical TOPS which can only ever execute within the 32MB SRAM that Tesla has created. Otherwise, Tesla's compute chip is stuck at 68 GBps LPDDR4 RAM. Pretty slow.
Pegasus uses HBM2 chips at 500GBps. Pegasus will be able to efficiently compute neural networks that are larger than 32MB in size.
Tesla is making big bets about this tiny 32MB SRAM. Bits and pieces of the CNN can fit in there, but almost certainly not the entire neural network.
You're right that this is a specialized chip. But even for NN / Deep Learning inference, it seems a bit underpowered to me from a RAM perspective