Earlier quoted context omitted.
Obviously, that's my point. We can do the math. GPT-4o can emit about 70 tokens a second. API pricing is $10/million for output tokens and $2.5/million for input tokens. Assuming a workload where inputs tokens are 10:1 with output tokens, and that I can generate continuous load (constantly generating tokens). I'll end up paying $210/day in API fees, or $76,650 in a year. Let's assume the hardware required to service…
That doesn't sounds like brilliant margins, to be honest. You've left out the entire "running a business" costs, plus the model training costs. They need to pay their staff, offices, and especially lawyers (for all the lawsuits over the scraped content used to train the models). It's not unusual for a startup to not be profitable, and they're obviously not as the company doesn't make a profit , but I'm not sure why i…
I was responding to someone upthread suggesting that they were running even inference at a loss.