Earlier quoted context omitted.
The GPUs that hosting providers are buying need to support vGPU features, which are only support by the higher-end workstation and server products.
Which is a software/binning issue (probably 100% software) not inherently HW because M60==GTX 980==GM204 and M40==GTX Titan X==Quadro M6000==GM200. Interestingly, for the first time ever, GP100 is unique (well, OK, K80 too but K80 was too late). And since Quadro P6000==Titan XP==GP102, it's probably just a software block here as well. Also, for the first time ever, the high-end Quadro will be the best FP32 GPU of it'…
To overcome this you need to disable UEFI boot/GPU boot in the BIOS and blacklist the card in the host OS and then create a PCI-stub device that will be used for the passthrough.
This is an utter and complete hack and you can't really use for production grade implementations.
I'm not sure that Quadro actually supports vGPU also AFAIK only Tesla and Grid parts do, Tesla does have some additional in GPU support for virtualization, i don't know how GRID handles it. GRID parts are less for compute and more for game/video streaming so they aren't fully virtualized, AFAIK they do not support P2P gpu communication or shared memory access, they are only more Hypervisor friendly so they can be thin provisioned.