Earlier quoted context omitted.
Wouldn't improving LLM efficiency make them even more useful across the board, then they can enjoy the nice economies of scale? The plan is to have LLM working completely autonomously, in that case, the more resources you have, the better. Perhaps people will use local LLM to ask questions, or coders use them for their personal projects, but that's not where the real money is.
If apple puts an inference SOC in their phone, the datacenters are all dead.
Apple's own desktops, with the fastest Apple Silicon GPUs and TDPs 20x higher than an iPhone still can't compete for real-world datacenter use even with RDNA clustering. Apple's GPGPU architecture is behind AMD at this point, there's a reason why Apple Intelligence is critically reliant on Nvidia and Google to provide inference backends.