Earlier quoted context omitted.
I believe this is a CPU/GPU vs ASIC comparison, rather than CPU vs GPU. They have always(ish) coexisted, being optimized for different things: ASICs have cost/speed/power advantages, but the design is more difficult than writing a computer program, and you can't reprogram them. Generally, you use an ASIC to perform a specific task. In this case, I think the takeaway is the LLM functionality here is performance-sensit…
It reminds me of the switch from GPUs to ASICs in bitcoin mining. I've been expecting this to happen.
AI being static weights is already challenged with the frequent model updates we already see - but may even be a relic once we find a new architecture.