Live data from Hacker News

No, it doesn't cost Anthropic $5k per Claude Code user

martinalderson.com

181–190 of 374 posts

Re: No, it doesn't cost Anthropic $5k per Claude Code user

#181

Earlier quoted context omitted.

Agree, but I guess the Opus 4.6 is 10x larger, rather than Chinese models being 10x more efficient. It is said that GPT-4 is already a 1.6T model, and Llama 4 behemoth is also much bigger than Chinese open-weight models. Chinese tech companies are short of frontier GPUs, but they did a lot of innovations on inference efficiency (Deepseek CEO Liang himself shows up in the author list of the related published papers).

No, Opus cannot be 10x larger than the chinese models. If Opus was 10x larger than the chinese models, then Google Vertex/Amazon Bedrock would serve it 10x slower than Deepseek/Kimi/etc. That's not the case. They're in the same order of magnitude of speed.

They serve it about 2x slower. So it must have about 2x the active parameters.

It could still be 10x larger overall, though that would not make it 10x more expensive.

Re: No, it doesn't cost Anthropic $5k per Claude Code user

#182

A huge number of people are convinced that OpenAI and Anthropic are selling inference tokens at a loss despite the fact that there's no evidence this is true and a lot of evidence that it isn't. It's just become a meme uncritically regurgitated. This sloppy Forbes article has polluted the epistemic environment because now theres a source to point to as "evidence." So yes this post author's estimation isn't perfect bu…

I'd love to be a fly on the wall when this argument is tried in front of a bankruptcy court. It drives me nuts. Of course there's evidence that they're selling tokens at a loss. The only thing these companies sell are tokens. That's their entire output. OpenAI is trying to build an ad business but it must be quite small still relative to selling tokens because I've not yet seen a single ad on ChatGPT. It's not like t…

The article is about compute cost though. By "lose money on inference" I mean the assertion that inference has negative gross margins which a lot of people truly believe. This is important because it's common to reason from this that LLM's are uneconomical and a ticking time bomb where prices will have to be jacked up several orders of magnitude just to cover the compute used for the tokens.

Re: No, it doesn't cost Anthropic $5k per Claude Code user

#183

Earlier quoted context omitted.

>It is not. It's a terrible comparison. Qwen, deepseek and other Chinese models are known for their 10x or even better efficiency compared to Anthropic's. I find it a good comparison because it is a good baseline since we have zero insider knowledge of Anthropic. They give me an idea that a certain size of a model has a certain cost associated. I don't buy the 10x efficiency thing: they are just lagging behind the pe…

> I don't buy the 10x efficiency thing: they are just lagging behind the performance of current SOTA models. They perform much worse than the current models while also costing much less - exactly what I would expect. Define "much worse". +--------------------------------------+-------------+-----------+------------------+ | Benchmark | Claude Opus | DeepSeek | DeepSeek vs Opus | +-------------------------------------…

Everyone who's used Opus knows it's better than the others in a way that isn't captured by the benchmarks. I would describe it as taste.

Lots of models get really close on benchmarks, but benchmarks only tell us how good they are at solving a defined problem. Opus is far better at solving ill-defined ones.

Re: No, it doesn't cost Anthropic $5k per Claude Code user

#184
post #99

Earlier quoted context omitted.

> That's a tautology. People think chinese models are 10x more efficient because they're 10x cheaper They do have different infrastructure / electricity costs and they might not run on nvidia hardware. It's not just the models.

Except there are providers that serve both chinese models AND opus as well. On the same hardware. Namely, Amazon Bedrock and Google Vertex. That means normalized infrastructure costs, normalized electricity costs, and normalized hardware performance. Normalized inference software stack, even (most likely). It's about a close of a 1 to 1 comparison as you can get. Both Amazon and Google serve Opus at roughly ~1/2 the…

And Microsoft's Azure. It's on all 3 major cloud providers. Which tells me, they can make profit from these cloud providers without having to pay for any hardware. They just take a small enough cut.

https://code.claude.com/docs/en/microsoft-foundry

https://www.anthropic.com/news/claude-in-microsoft-foundry

Re: No, it doesn't cost Anthropic $5k per Claude Code user

#186
post #179
post #175

Earlier quoted context omitted.

> unless China is pumping out domestic chips cheaply enough They are. Nvidia makes A LOT of profit. Hey, top stock for a reason. > I wouldn't be surprised if things like DS were trained and now hosted on Nvidia hardware DS is "old". I wouldn't study them. The new 1s have a mandate to at least run on local hardware. There are data center requirements. I agree it could still be trained on Nvidia GPUs (black market etc)…

> The new 1s have a mandate to at least run on local hardware. They do? Source? But if that's true, it would explain why Minimax, Z.ai and Moonshot are all organized as Singaporean holding companies, with claimed data center locations (according to OpenRouter) in the US or Singapore and only the devs in China. Can't be forced to use inferior local hardware if you're just a body shop for a "foreign" AI company. ;)

> with claimed data center locations (according to OpenRouter) in the US or Singapore and only the devs in China

They just have a China only endpoint and likely a company under a different name.

Nothing to do with AI. TikTok is similar (global vs China operations).

Re: No, it doesn't cost Anthropic $5k per Claude Code user

#187

Earlier quoted context omitted.

> I don't buy the 10x efficiency thing: they are just lagging behind the performance of current SOTA models. They perform much worse than the current models while also costing much less - exactly what I would expect. Define "much worse". +--------------------------------------+-------------+-----------+------------------+ | Benchmark | Claude Opus | DeepSeek | DeepSeek vs Opus | +-------------------------------------…

Everyone who's used Opus knows it's better than the others in a way that isn't captured by the benchmarks. I would describe it as taste. Lots of models get really close on benchmarks, but benchmarks only tell us how good they are at solving a defined problem. Opus is far better at solving ill-defined ones.

>Everyone who's used Opus knows it's better than the others in a way that isn't captured by the benchmarks. I would describe it as taste.

Ah, the "trust me bro" advantage. Couldn't it just be brand identity and familiarity?

Re: No, it doesn't cost Anthropic $5k per Claude Code user

#188
post #99

Earlier quoted context omitted.

> That's a tautology. People think chinese models are 10x more efficient because they're 10x cheaper They do have different infrastructure / electricity costs and they might not run on nvidia hardware. It's not just the models.

Except there are providers that serve both chinese models AND opus as well. On the same hardware. Namely, Amazon Bedrock and Google Vertex. That means normalized infrastructure costs, normalized electricity costs, and normalized hardware performance. Normalized inference software stack, even (most likely). It's about a close of a 1 to 1 comparison as you can get. Both Amazon and Google serve Opus at roughly ~1/2 the…

> Both Amazon and Google serve Opus at roughly ~1/2 the speed of the chinese models

We were responded about 10x not 0.5x.

x86 vs arm64 could have different performance. The Chinese models could be optimized for different hardware so it could show massive differences.

Re: No, it doesn't cost Anthropic $5k per Claude Code user

#189
post #106

I calculated only last weekend that my team would cost, if we would run Claude Code on retail API costs, around $200k/mo. We pay $1400/month in Max subscriptions. So that's $50k/user... But what tokens CC is reporting in their json -> a lot of this must be cached etc, so doubt it's anywhere near $50k cost, but not sure how to figure out what it would cost and I'm sure as hell not going to try.

I'm surprised, isn't it forbidden to use the Max plan as part of a company? Just curious, as I thought it was forbidden by the ToS but I'm not sure if I have a good understanding of it

Most companies forbid it though, since you're not covered by any legal protection - for example, Anthropic can use your data or code to train new models and more.

Re: No, it doesn't cost Anthropic $5k per Claude Code user

#190
I think the main issue I have with the article is that author whole argument is based on 'Qwen wouldn't run at a loss'. But why wouldn't it? Depsite it being a business, there might be a number of arguments why they decide to run without profit for now: from trying to expand the user base, to Chinese government sponsoring Chinese AI business.
Post reply on HN