Earlier quoted context omitted.
This is the most script kiddy comment I've seen in a while. llama.cpp is just inference, not training, and the CUDA backend is still the fastest one by far. No one is even close to matching CUDA on either training or inference. The closest is AMD with ROCm, but there's likely a decade of work to be done to be competitive.
Yes, and inference is a huge market in itself and potentially larger than training (gut feeling haven’t run numbers) Keep NVIDIA for training and Intel/AMD/Cerebras/… for interference.
Inference is also a much smaller market right now, but will likely be overtaken later as we have more people using the models than competing to train the best one.