> Realistically, serving 300 users per GPU you'll spend a lifetime cost of about $133 per user, plus the datacenter/upkeep bill. What is the operational cost and when does it become more expensive than the upfront capex? The B200 tops out at 1000W and idles around 140W. It averages around 600W. https://www.lightly.ai/blog/nvidia-b200-vs-h100 U.S. average electricity cost is $.14 per kWh in March. https://www.eia.gov/…
I wonder what the power costs are when you put jet turbines in front of your DC to power it.
Inference cost at scale with napkin math
11–20 of 20 posts
Re: Inference cost at scale with napkin math
#12what kind of math is this? why isn't it B = 562 / 2 = 281?
Re: Inference cost at scale with napkin math
#13> We'll assume a 32B dense model, as they've have gotten quite good for production use and a B200 can comfortably serve them. This could be a Gemma, Qwen, DeepSeek, whatever. That seems like a very consequential point to include halfway through the post. They aren't wrong that Qwen 3.6 26B or Gemma 4 31B are quite good, depending on the use case, but if we're doing napkin math, I'd want some more headroom in the assu…
But in reality, 32B dense is very similar* to 32B activated on MoE in terms of inference costs. And I highly suspect eg Opus is around that level of active params.
A 284ba13b model at scale, is almost certainly cheaper to serve than a 32b dense model.
*as you can shard the model across multiple GPUs at scale. but in reality you have some loss of efficiency from GPU coordination and expert routing
Re: Inference cost at scale with napkin math
#14> We'll assume a 32B dense model, as they've have gotten quite good for production use and a B200 can comfortably serve them. This could be a Gemma, Qwen, DeepSeek, whatever. That seems like a very consequential point to include halfway through the post. They aren't wrong that Qwen 3.6 26B or Gemma 4 31B are quite good, depending on the use case, but if we're doing napkin math, I'd want some more headroom in the assu…
Yes 32B dense is a weird one to choose. But in reality, 32B dense is very similar* to 32B activated on MoE in terms of inference costs. And I highly suspect eg Opus is around that level of active params. A 284ba13b model at scale, is almost certainly cheaper to serve than a 32b dense model. *as you can shard the model across multiple GPUs at scale. but in reality you have some loss of efficiency from GPU coordination…
Re: Inference cost at scale with napkin math
#15Re: Inference cost at scale with napkin math
#16I'd like to see a bit of the running costs inside the napkin math. Power, cooling, maintenance, rent, etc. are probably significant factors as well.
Re: Inference cost at scale with napkin math
#17> Realistically, serving 300 users per GPU you'll spend a lifetime cost of about $133 per user, plus the datacenter/upkeep bill. What is the operational cost and when does it become more expensive than the upfront capex? The B200 tops out at 1000W and idles around 140W. It averages around 600W. https://www.lightly.ai/blog/nvidia-b200-vs-h100 U.S. average electricity cost is $.14 per kWh in March. https://www.eia.gov/…
Re: Inference cost at scale with napkin math
#18>This largely depends on whether you own or rent your hardware. At $40,000 per B200, your lifetime cost per user is 40_000/num_users. In the 100% duty cycle case (worst for cost), that's 6k$ per user. Realistically, serving 300 users per GPU you'll spend a lifetime cost of about $133 per user, plus the datacenter/upkeep bill . If you rent the GPU, the cost is more straightforward. At an hourly rate of $43, your hourl…
This cannot be done on most premises because of power, noise, and cooling.
Re: Inference cost at scale with napkin math
#19Earlier quoted context omitted.
Yes 32B dense is a weird one to choose. But in reality, 32B dense is very similar* to 32B activated on MoE in terms of inference costs. And I highly suspect eg Opus is around that level of active params. A 284ba13b model at scale, is almost certainly cheaper to serve than a 32b dense model. *as you can shard the model across multiple GPUs at scale. but in reality you have some loss of efficiency from GPU coordination…
That's good information. I couldn't possibly even start to run even DeepSeek Flash on my system, but also if you're assuming multiple GPUs, that is going to affect the napkin math.
Re: Inference cost at scale with napkin math
#20Earlier quoted context omitted.
So what's the cost separating them from placing this box at their premise? Network throughout?
Plus power and cooling.
Not having physical access to my assets doesn't sound secure at all, and even a residential internet connection could handle this throughout.