Earlier quoted context omitted.
You use at least half of this stack for desktop setups. You need copying daemons, the ecosystem support (docker-nvidia, etc.), some of the libraries, etc. even when you're on a single system. If you're doing inference on a server; MIG comes into play. If you're doing inference on a larger cloud, GPU-direct storage comes into play. It's all modular.
No you don‘t need much bandwidth between cards for inference
GPU-Direct is about pumping data from storage devices to cards, esp. from high speed storage systems across networks.
MIG actually shares a single card to multiple instances, so many processes or VMs can use a single card for smaller tasks.
Nothing I have written in my previous comment is related to inter-card, inter-server communication, but all are related to disk-GPU, CPU-GPU or RAM-CPU communication.
Edit: I mean, it's not OK to talk about downvoting, and downvote as you like but, I install and enable these cards for researchers. I know what I'm installing and what it does. C'mon now. :D