Live data from Hacker News

Are OpenAI and Anthropic losing money on inference?

martinalderson.com

441–450 of 495 posts

Re: Are OpenAI and Anthropic losing money on inference?

#441

Earlier quoted context omitted.

Why wouldn't you factor in training? It is not like you can train once and then have the model run for years. You need to constantly improve to keep up with the competition. The lifespan of a model is just a few months at this point.

I suspect we've already reached the point with models at the GPT5 tier where the average person will no longer recognize improvements and this model can be slightly improved at slow intervals and indeed run for years. Meanwhile research grade models will still need to be trained at massive cost to improve performance on relatively short time scales.

The "Pro" variant of GTP-5 is probably the best model around and most people are not even aware that it exists. One reason is that as models get more capable, they also get a lot more expensive to run so this "Pro" is only available at the $200/month pro plan.

At the same time, more capable models are also a lot more expensive to train.

The key point is that the relationship between all these magnitudes is not linear, so the economics of the whole thing start to look wobbly.

Soon we will probably arrive at a point where these huge training runs must stop, because the performance improvement does not match the huge cost increase, and because the resulting model would be so expensive to run that the market for it would be too small.

Re: Are OpenAI and Anthropic losing money on inference?

#442

Ok, one issue I have with this analysis is the breakdown between input and output tokens. I'm the kind of person who spend most of my chat asking questions, so I might only use 20ish input tokens per prompt, where Gemini is having to put out several hundred, which would seem to affect the economics quite a bit

Yeah, I've noticed Chatgpt5 is very chatty. I can ask a 1 sentence question and get back 3-4 paragraphs, most of which I ignore, depending upon the task.

I haven’t used it without customization, but I find it follows my brevity user instructions more strictly.

Re: Are OpenAI and Anthropic losing money on inference?

#443
post #247

Earlier quoted context omitted.

From the latest NYT Hard Fork podcast [1]. The hosts were invited to a dinner hosted by Sam, where Sam said "we're profitable if we remove training from the equation", they report he turned to Lightcap (COO) and asked "right?" and Lightcap gave an "eeekk we're close". They aren't yet profitable even just on inference, and its possible Sam didn't know that until very recently. [1] https://www.nytimes.com/2025/08/22/po…

“We’re not profitable even if we discount training costs.” and “Inference revenue significantly exceeds inference costs.” are not incompatible statements. So maybe only the first part of Sam’s comment was correct.

I should have provided a direct quote:

> At first, he answered, no, we would be profitable, if not for training new models. Essentially, if you take away all the stuff, all the money we’re spending on building new models and just look at the cost of serving the existing models, we are profitable on that basis. And then he looked at Brad Lightcap, who is the COO. And he sort of said, right? And Brad kind of squirmed in his seat a little bit and was like, well — He’s like, we’re pretty close.

I don't think you can square that with what he stated in the Axios article:

> "We're profitable on inference. If we didn't pay for training, we'd be a very profitable company."

Except, if the NYT dinner happened after the Axios article interview, which is possible given when each was published, and he was actually literally unaware of the company's financials.

Personally: it feels like it should reflect very poorly on OpenAI that their CEO has been, charitably, entirely unaware how close they are to profitability (and uncharitably, that he actively lies about it). But I'm not sure if the broader news cycle caught it; the only place I've heard this mentioned is literally this NYT Hard Fork podcast which is hosted by the people who were at the dinner.

Re: Are OpenAI and Anthropic losing money on inference?

#444

Earlier quoted context omitted.

>>This is clearly the case ... probably >>OpenAI would be incredibly profitable if they could live off the profit from improved inference efficiency on a GPT4 level model! If gpt4 was basically free money at this point it's real weird that their first instinct was to cut it off after gpt5

> If gpt4 was basically free money at this point it's real weird that their first instinct was to cut it off after gpt5 People find the UX of choosing a model very confusing, the idea with 5 is that it would route things appropriately and so eliminate this confusion. That was the motivation for removing 4. But people were upset enough that they decided to bring it back for a while, at least.

They picked the worst possible time to make the change if money wasn’t involved (which is why I assumed GPT-5 must be massively cheaper to run). The backlash from being forced to use it cost a fair bit of the model’s reputation.

Re: Are OpenAI and Anthropic losing money on inference?

#445
post #407

This article's math is wrong on many fundamental levels. One of the most obvious ones is that prefill is nowhere near bandwidth bound. If you compute out the MFU the author gets it's 1.44 million input tokens per second * 37 billion active params * 2 (FMA) / 8 [GPUs per instance] = 13 Petaflops per second. That's approximately 7x absolutely peak FLOPS on the hardware. Obviously, that's impossible. There's many other…

Agree that the writeup is very wrong, especially for the output tokens. Here is how anyone with enough money to allocate a small cluster of powerful GPUs can decode huge models at scale, since nearly 4 months ago, with costs of 0.2 USD/million output tokens. https://lmsys.org/blog/2025-05-05-large-scale-ep/ This has gotten significantly cheaper yet with additional code hacks since then, and with using the B200s.

You can also look at the price of opensource models on openrouter, which are a fraction of the cost of closed source models. This is a market that is heavily commoditized, so I would expect it reflect the true cost with a small margin.

Re: Are OpenAI and Anthropic losing money on inference?

#446

Earlier quoted context omitted.

Why wouldn't you factor in training? It is not like you can train once and then have the model run for years. You need to constantly improve to keep up with the competition. The lifespan of a model is just a few months at this point.

I suspect we've already reached the point with models at the GPT5 tier where the average person will no longer recognize improvements and this model can be slightly improved at slow intervals and indeed run for years. Meanwhile research grade models will still need to be trained at massive cost to improve performance on relatively short time scales.

I may not qualify as an "average user" but I shudder imagining being stuck using a 1+ yr stale model for development given my experiences using a newer framework than what was available during training.

Passing in docs usually helps, but I've had some incredibly aggravating experiences where a model just absolutely cannot accept their "mental mode" is incorrect and that they need to forget the tens of thousands of lines of out of date example code they've ingested during training. IMO it's an under-discussed aspect of the current effectiveness of LLM development thanks to the training arms race.

I recently had to fight Gemini to accept that a library (a Google developed AI library for JS, somewhat ironically) had just released a major version update with a lot of API changes that invalidated 99% of the docs and example code online. And boy was there a lot of old code floating around thanks to the vast amounts of SEO blog spam for anything AI adjacent.

Re: Are OpenAI and Anthropic losing money on inference?

#447

Earlier quoted context omitted.

Anyone paying attention should have zero trust in what Sam Altman says.

What do you think his strategy is? He has to make money at some point. I don’t buy the logic that he will “scam” his investors and run away at some point.

> He has to make money at some point.

Yes, but two paths to doing that are to a) build a profitable company, and b) accumulate personal wealth and walk away from a non-profitable company.

I'm not saying OpenAI is unprofitable, but nor do I see Altman as the sort who'd rule out option b.

Re: Are OpenAI and Anthropic losing money on inference?

#448

Earlier quoted context omitted.

> If gpt4 was basically free money at this point it's real weird that their first instinct was to cut it off after gpt5 People find the UX of choosing a model very confusing, the idea with 5 is that it would route things appropriately and so eliminate this confusion. That was the motivation for removing 4. But people were upset enough that they decided to bring it back for a while, at least.

They picked the worst possible time to make the change if money wasn’t involved (which is why I assumed GPT-5 must be massively cheaper to run). The backlash from being forced to use it cost a fair bit of the model’s reputation.

Yeah it didnt work out for them, for sure.

Re: Are OpenAI and Anthropic losing money on inference?

#449
post #371

Earlier quoted context omitted.

I wonder how much capex risk there is in this model, depreciating the GPUs over 5 years is fine if you can guarantee utilization. Losing market share might be a death sentence for some of these firms as utilization falls.

What I hear nobody talking about is the price elasticity of demand and how this plays into the economics of the model business.

I think some of the power user demand is fairly inelastic. I’ve seen developers who are allergic to spending money happily drop $200/mo on those new Claude subscriptions.

Re: Are OpenAI and Anthropic losing money on inference?

#450

Earlier quoted context omitted.

In the same way that every other startup tries to sweep R&D costs under the rug and say “yeah but the marginal unit economics have 50% gross margins, we’ll be a great business soon”.

lol. TBH I don't take anyone seriously unless they are talking about cash flows (FCFF or FCFE specifically). Who cares about expense classification - show me the money!

Google and Facebook had negative free cash flow for years early in their lives. All the good investors were lolling at the bad investors lolling at the cash they were burning.
Post reply on HN