Earlier quoted context omitted.
It is beyond porting, it is mentality of developers. Change is expensive and I'm not just talking about $ value. > Also, Google has been using TPUs for a long time now, and __they__ never hit a brick wall for a lack of CUDA. That's exactly what I'm saying. __they__ is the keyword.
Not sure what you mean. Google is a big company. Their TPUs have many users internally.
If you're going to design a custom chip and deploy it in your data centers, you're also committing to hiring and training developers to build for it.
That's a kind of moat, but with private chips. While you solve one problem (getting the compute you want), you create another: supporting and maintaining that ecosystem long term.
NVIDIA was successful because they got their hardware into developers hands, which created a feedback loop, developers asked for fixes/features, NVIDIA built them, the software stack improved, and the hardware evolved alongside it. That developer flywheel is what made CUDA dominant and is extremely hard to replicate because the shortage of talented developers is real.