Earlier quoted context omitted.
I think it's niche now because getting the hardware to run it is expensive and the quantized models don't work as well. If those improve then it would be a no brainer to pay one off for the hardware instead of a fortune for API calls.
AI vendors are attempting to offer the whole apple. And they are spending huge sums of money in the process. But most businesses don't really care about most of the apple --- they only need their special bite out of it. For example, doctors mainly care about medicine. Nvidia is attempting to provide the hardware needed for local, specialized models.
But I don’t know about specialised: this could run quite large models with MoE.