This will be interesting for a few reasons. First, depending on where the median pricing settles w/ 3rd party providers will tell us what it costs to serve a 3T model. Since it's going to be mxfp4 native, it'll take ~1.5TB of VRAM to host this, which is juuust at the limit of 8xb200s (but realistically you'll need 16x for context / throughput optimisation). Won't be cheap to host, but at least we should get some rang…
You're assuming inference providers are going to sell tokens at cost. You're also assuming that the inference providers have will optimized inference engine. I haven't seen that to be the case so far, to be honest. Take a look at GLM 5 vs GLM 5.2 pricing -- GLM 5.2 cost more despite being the same model. Take a look a look at DeepSeek, which hosts DS v4, profitably, yet others aren't able or willing to match the pric…
OTOH, the multiple providers who have settled around the same price point ($3.48/M output tokens for multiple providers with good reputations) does indicate where it is profitable: https://openrouter.ai/deepseek/deepseek-v4-pro#providers