Claude subscription is equivalant of spot instance And APIs are on-demand service equivalant. Priority is set to APIs and leftover compute is used by Subscription Plans. When there is no capacity, subscriptions are routed to Highly Quantized cheaper models behind the scenes. Selling subscription makes it cheaper to run such inference at scale otherwise many times your capacity is just sitting there idle. Also, these…
> Claude is 2x better than Codex This hasn't been true in a long time.
No, it doesn't cost Anthropic $5k per Claude Code user
121–130 of 374 posts
Re: No, it doesn't cost Anthropic $5k per Claude Code user
#122Earlier quoted context omitted.
Agree, but I guess the Opus 4.6 is 10x larger, rather than Chinese models being 10x more efficient. It is said that GPT-4 is already a 1.6T model, and Llama 4 behemoth is also much bigger than Chinese open-weight models. Chinese tech companies are short of frontier GPUs, but they did a lot of innovations on inference efficiency (Deepseek CEO Liang himself shows up in the author list of the related published papers).
No, Opus cannot be 10x larger than the chinese models. If Opus was 10x larger than the chinese models, then Google Vertex/Amazon Bedrock would serve it 10x slower than Deepseek/Kimi/etc. That's not the case. They're in the same order of magnitude of speed.
According to OpenRouter, AWS serves the latest Opus and Sonnet at roughly the same speed. It's likely that they simply allocate hardware differently per model.
Re: No, it doesn't cost Anthropic $5k per Claude Code user
#123This is such a well-written essay. Every line revealed the answer to the immediate question I had just thought of
I can’t get past all the LLM-isms. Do people really not care about AI-slopifying their writing? It’s like learning about bad kerning, you see it everywhere.
Re: No, it doesn't cost Anthropic $5k per Claude Code user
#124Claude subscription is equivalant of spot instance And APIs are on-demand service equivalant. Priority is set to APIs and leftover compute is used by Subscription Plans. When there is no capacity, subscriptions are routed to Highly Quantized cheaper models behind the scenes. Selling subscription makes it cheaper to run such inference at scale otherwise many times your capacity is just sitting there idle. Also, these…
> Claude is 2x better than Codex This hasn't been true in a long time.
Maybe that's just CLAUDE.md and memory causing the difference of course.
As a matter of preference however I like the way Claude Code works just a lot better, instructing it to work with parallel subagents in work trees etc. just matches the way I think these things should work I guess.
Re: No, it doesn't cost Anthropic $5k per Claude Code user
#125Earlier quoted context omitted.
But opportunity cost is not actual cost. “If everyone just kept paying but used our service less we would be more profitable” is true, but not in any meaningful way. Are Anthropic currently unable to sell subscriptions because they don’t have capacity?
> Are Anthropic currently unable to sell subscriptions because they don’t have capacity? Absolutely! Im currently paying $170 to google to use Opus in antigravity without limit in full agent mode, because I tried Anthropic $20 subscription and busted my limit within a single prompt. Im not gonna pay them $200 only to find out I hit the limit after 20 or even 50 prompts. And after 2 more months my price is going to do…
Re: No, it doesn't cost Anthropic $5k per Claude Code user
#126What people don't realize is that cache is *free*, well not free, but compared to the compute required to recompute it? Relatively free. If you remove the cached token cost from pricing the overall api usage drops from around $5000 to $800 (or $200 per week) on the $200 max subscription. Still 4x cheaper over API, but not costing money either - if I had to guess it's break even as the compute is most likely going idl…
I'm incredibly salty about this - they're essentially monetizing intensely something that allows them to sell their inference at premium prices to more users - without any caching, they'd have much less capacity available.
Re: No, it doesn't cost Anthropic $5k per Claude Code user
#127Earlier quoted context omitted.
Try 10s of trillions. These days everyone is running 4-bit at inference (the flagship feature of Blackwell+), with the big flagship models running on recently installed Nvidia 72gpu rubin clusters (and equivalent-ish world size for those rented Ironwood TPUs Anthropic also uses). Let's see, Vera Rubin racks come standard with 20 TB (Blackwell NVL72 with 10 TB) of unified memory, and NVFP4 fits 2 parameters per btye..…
Nobody is running 10s of trillion param models in 2026. That's ridiculous. Opus is 2T-3T in size at most .
Re: No, it doesn't cost Anthropic $5k per Claude Code user
#128Claude subscription is equivalant of spot instance And APIs are on-demand service equivalant. Priority is set to APIs and leftover compute is used by Subscription Plans. When there is no capacity, subscriptions are routed to Highly Quantized cheaper models behind the scenes. Selling subscription makes it cheaper to run such inference at scale otherwise many times your capacity is just sitting there idle. Also, these…
Have they announced this?
Re: No, it doesn't cost Anthropic $5k per Claude Code user
#129If Anthropic's compute is fully saturated then the Claude code power users do represent an opportunity cost to Anthropic much closer to $5,000 then $500. Anthropic's models may be similar in parameter size to model's on open router, but none of the others are in the headlines nearly as much (especially recently) so the comparison is extremely flawed. The argument in this article is like comparing the cost of a Rolex…
Re: No, it doesn't cost Anthropic $5k per Claude Code user
#130Earlier quoted context omitted.
I'm fascinated to know the kind of work that allows you to intelligently allocate so much resources. I use Claude extensively and feel that I great value out of it but I reach a limit in terms of what I can do that makes sense relatively quickly it seems.
[flagged]
I wanted to believe that you're essentially trolling, but no - that service exist. And not an upstart, there is coverage going back several years.
Our societies are seriously fucked.