Earlier quoted context omitted.
What does performance look like?
Spoiler alert: not good enough to break CUDA's moat
Inference side is partly about performance, but mostly about cost per token.
And given that there has been a ton of standardization around LLaMA architectures, AMD/ROCm can target this much more easily, and still take a nice chunk of the inference market for non-SOTA models.