If you're doing a lot of model training, buying GPUs or long-term reservations of GPUs is a no-brainer. But when it comes to inference, latency matters and it gets trickier talking between e.g. your AWS infra and your GPUs somewhere else.
It seems lots of providers can give you enough to get by doing inference in a company's earliest stages. But what if I need hundreds or thousands of A100s during peak usage? Is anyone doing this successfully with a non-hyperscaler?