Do people think that nobody at nVidia has ever heard of specialized deep learning processors? 1. Volta GPUs already have little matmul cores, basically a bunch of little TPUs. 2. The graphics dedicated silicon is an extremely tiny portion of the die, a trivial component (source: Bill Dally, nVidia chief scientist). 3. Memory access power and performance is the bottleneck (even in the TPU paper), and will only continu…
All that said, Bill Daly rocks, and NVDA is a hardened target. But the DL frameworks have enormous performance holes once one stops running Resnet-152 and other popular benchmark graphs in the same way that 3DMark performance is not necessarily representative of actual gaming performance unless NVDA took it upon themselves to make it so.
And since DL is such a dynamic field (just like game engines), I expect this situation to persist for a very, very long time.