This is a game changer for local inference IMO, where you will basically never need to do this. On the other hand in a serverless/cloud setting it may be more problematic if you don't want to strand GPUs with locally attached HBF.