If they never go public, there's our answer as well.
No, it doesn't cost Anthropic $5k per Claude Code user
291–300 of 374 posts
Re: No, it doesn't cost Anthropic $5k per Claude Code user
#292Earlier quoted context omitted.
Everyone who's used Opus knows it's better than the others in a way that isn't captured by the benchmarks. I would describe it as taste. Lots of models get really close on benchmarks, but benchmarks only tell us how good they are at solving a defined problem. Opus is far better at solving ill-defined ones.
One of the main edges Anthropic has is that "personality tuning" gap. "Nice to use" is a differentiator when raw performance isn't. OpenAI can sometimes get an edge over Anthropic in hard narrow STEM tasks. I trust benchmarks over vibes there - and the benchmarks show the teams trading blows release after release. Tracking Claude Code vs OpenAI Codex on SWE-bench Verified feels like watching the back alley knife figh…
Re: No, it doesn't cost Anthropic $5k per Claude Code user
#293Earlier quoted context omitted.
Try 10s of trillions. These days everyone is running 4-bit at inference (the flagship feature of Blackwell+), with the big flagship models running on recently installed Nvidia 72gpu rubin clusters (and equivalent-ish world size for those rented Ironwood TPUs Anthropic also uses). Let's see, Vera Rubin racks come standard with 20 TB (Blackwell NVL72 with 10 TB) of unified memory, and NVFP4 fits 2 parameters per btye..…
Nobody is running 10s of trillion param models in 2026. That's ridiculous. Opus is 2T-3T in size at most .
You have presented a vibe-based rebuttal with no evidence or or logic to outline why you think labs are still stuck in the single trillions of parameters (GPT 4 was ~1 trillion params!). Though, you have successfully cunninghammed me into saying that while anything I publicly state is derived from public info, working in the industry itself is a helpful guide to point at the right public info to reference.
Re: No, it doesn't cost Anthropic $5k per Claude Code user
#294"Any conversation about token costs devolves into an ad-hoc, informally-specified, bug-ridden implementation of half of generally accepted accounting principles." We have a way of determining if Anthropic is, or has the capability of being profitable, and what the levers to that may be. AI may be world-changing, but the accounting principles behind AI labs are no different than those behind a Pizza Hut. Even if the c…
Re: No, it doesn't cost Anthropic $5k per Claude Code user
#295"Any conversation about token costs devolves into an ad-hoc, informally-specified, bug-ridden implementation of half of generally accepted accounting principles." We have a way of determining if Anthropic is, or has the capability of being profitable, and what the levers to that may be. AI may be world-changing, but the accounting principles behind AI labs are no different than those behind a Pizza Hut. Even if the c…
However, the GAAP P&L tells the opposite story. You book $200M revenue in the same year you spend $1B training the next model, so you report an $800M loss. Next year you book $2B against $10B in training spend, reporting an $8B loss. The business looks like it's dying when every individual model generation actually generates a healthy profit.
That's actually Dario's answer to your depreciation question. If each cohort earns back its training cost within its natural lifespan (however short that lifespan is), the depreciation schedule is already baked in. The model doesn't need to live forever, it just needs to return more than it cost before the next one replaces it. Whether that's actually happening at Anthropic is a different question, and one we can't answer without audited financials, but it's the claim Dario makes (and seems entirely reasonable from a distance).
Re: No, it doesn't cost Anthropic $5k per Claude Code user
#296Earlier quoted context omitted.
The article is about compute cost though. By "lose money on inference" I mean the assertion that inference has negative gross margins which a lot of people truly believe. This is important because it's common to reason from this that LLM's are uneconomical and a ticking time bomb where prices will have to be jacked up several orders of magnitude just to cover the compute used for the tokens.
But there's no such thing as compute cost in the abstract. What exactly is compute cost for AI? Does it include: • Inference used for training? Modern training pipelines aren't just gradient descent, there's a ton of inference used in them too. • Gradient descent itself? • The CPUs and disks storing and managing the datasets? • The web servers? • The people paid to swap out failed components at the dc? Let's say you…
The API price should hopefully incorporate the capitalized cost of the hardware, the facility rent, the cost to train the model, the r&d, cost of sales, etc., to make it profitable.
Claude Code Max may be able to offer a good price by having a mix of higher and lower utilization of users and ignoring the fixed costs, treating it as a driver of API sales. But it doesn't make sense to essentially pay people to use it.
Re: No, it doesn't cost Anthropic $5k per Claude Code user
#297"Any conversation about token costs devolves into an ad-hoc, informally-specified, bug-ridden implementation of half of generally accepted accounting principles." We have a way of determining if Anthropic is, or has the capability of being profitable, and what the levers to that may be. AI may be world-changing, but the accounting principles behind AI labs are no different than those behind a Pizza Hut. Even if the c…
Maybe not? This is an argument that has to be made using numbers. We can't do the estimate without the numbers.
Re: No, it doesn't cost Anthropic $5k per Claude Code user
#298"Any conversation about token costs devolves into an ad-hoc, informally-specified, bug-ridden implementation of half of generally accepted accounting principles." We have a way of determining if Anthropic is, or has the capability of being profitable, and what the levers to that may be. AI may be world-changing, but the accounting principles behind AI labs are no different than those behind a Pizza Hut. Even if the c…
Dario has made a specific cohort argument here. His numbers (from various interviews) are: you train a model in 2023 for $100M, deploy it, and it earns $200M over its lifetime. Meanwhile you train the 2024 model for $1B, which goes on to earn $2B. Each vintage returns 2x on its training cost. However, the GAAP P&L tells the opposite story. You book $200M revenue in the same year you spend $1B training the next model,…
Re: No, it doesn't cost Anthropic $5k per Claude Code user
#299"Any conversation about token costs devolves into an ad-hoc, informally-specified, bug-ridden implementation of half of generally accepted accounting principles." We have a way of determining if Anthropic is, or has the capability of being profitable, and what the levers to that may be. AI may be world-changing, but the accounting principles behind AI labs are no different than those behind a Pizza Hut. Even if the c…
Dario has made a specific cohort argument here. His numbers (from various interviews) are: you train a model in 2023 for $100M, deploy it, and it earns $200M over its lifetime. Meanwhile you train the 2024 model for $1B, which goes on to earn $2B. Each vintage returns 2x on its training cost. However, the GAAP P&L tells the opposite story. You book $200M revenue in the same year you spend $1B training the next model,…
And I admit that I made that assertion from my gut without actually knowing if it's true or not.
Re: No, it doesn't cost Anthropic $5k per Claude Code user
#300Earlier quoted context omitted.
Except there are providers that serve both chinese models AND opus as well. On the same hardware. Namely, Amazon Bedrock and Google Vertex. That means normalized infrastructure costs, normalized electricity costs, and normalized hardware performance. Normalized inference software stack, even (most likely). It's about a close of a 1 to 1 comparison as you can get. Both Amazon and Google serve Opus at roughly ~1/2 the…
Deployments like bedrock have no where near SOTA operational efficiency, 1-2 OOM behind. The hardware is much closer, but pipeline, schedule, cache, recomposition, routing etc optimizations blow naive end to end architectures out of the water.