And they may continue to win because successful execution of an 80% product is worth far more than a 90+% powerpoint processor cough TenSilica et al. cough, or because this is such a huge potential market, it might actually go to a successful competitor. 2018 and beyond will be very interesting.
For while it's really desirable have the deep learning equivalent of x86 assembly language (CUDA) across a full stack from training to inference, in the end, IMO cost will be king. I'm not a big fan of $150K high-end servers filled with $5000 GPUs that can be bested with clever code on a $25K server fill with $1200 consumer GPUs. But I am a huge fan of charging what you can while you are unopposed. It's just that I think that state is temporary.