Earlier quoted context omitted.
The cost of local hardware is amortized if a whole team uses it instead of just 1 dev (GPUs are extremely underutilized if you launch just 1 generation stream). I'm not sure why everyone always assumes solo devs with Macs. We've just ordered a large datacenter-grade node for use by the whole dev team, and the calculations show that it's going to cost the same amount of money if we kept using AWS Bedrock (infosec reas…
> GPUs are extremely underutilized if you launch just 1 generation stream why is that? b/c the thing is waiting for the hoooman and idling? or some parallelizable interleaving steps? I have no intuition yet how this works under the hood.
Waiting for the hooman (or tool calls) won't help either, of course.