Earlier quoted context omitted.
> This idea that the subscriptions are subsidized is repeated over and over, but I've never seen any proof of this. It seems to be entirely based on the inferred API cost the subscription usage could give you, but there are a lot of assumptions needed for that to follow. My claude code environment shows me cost per token used in that session, according to API costs. It regularly exceeds $200. I pay $200 a month for m…
That's what they want to charge you. Not the actual cost. The actual cost is a gpu that's probably already paid off and about $2 of electricity
Why current LLM costs are not sustainable
131–140 of 216 posts
Re: Why current LLM costs are not sustainable
#132It's weird to see people claiming that model capabilities are plateauing. It wasn't until late last year that we even had strong coding models. Imagine if, less than a year after the first iPhone launched, people claimed that smartphone capabilities were "plateauing" because Apple hadn't yet launched a new phone. And it seems the issue is less than "models aren't getting better" than, "models are good enough to handl…
Re: Why current LLM costs are not sustainable
#133If the subscription is gutted by factor 2/5/10/20/65 to make it more profitable for Anthropic it will be harder for users to justify the subscription.
On the other hand 13,000 USD in API credits can go a very long way if used ergonomically. For instance using a max context length of 200k is multiplying your reach in comparison to 1m context.
Re: Why current LLM costs are not sustainable
#134Earlier quoted context omitted.
I've spend a week doing just that - I said at API pricing, $200/month currently seems adequate for 2-4 weeks of usage for me at work. $50 would be 10M input tokens, not tens of thousands.
> I said at API pricing, $200/month Well I saw $200/month and thought you were talking about a max plan, sorry. But I will say unless you're using that top end model extremely judiciously $200 for 2-4 weeks of work is similarly hard to believe (see the other poster breaking down their usage). What are you typically doing? Must be pretty hardcore stuff if you need to use the baddest available model. How many interacti…
It is not my experience that you need to do 'hardcore stuff' to require the use of a large model. The difference in productivity between babysitting Sonnet and trying to get the result into a good shape compared to using Opus 4.8 seems large to me.
At home, unfortunately I only have the stats from the official apps rather than granular ones, and it looks like the Claude Desktop app is buggy: it was showing 17M tokens total in the last 30 days, but even just clicking on a conversation in my side bar increased the counter to 19M. It's clearly not working.
Codex shows up to 900M tokens total/week.
Re: Why current LLM costs are not sustainable
#135Earlier quoted context omitted.
>3. We're massively overusing SOTA models. As long as you're on a subsidized subscription, you can use Claude Opus 4.8 high to write blog article meta descriptions. If you paid by token, you wouldn't do that. This idea that the subscriptions are subsidized is repeated over and over, but I've never seen any proof of this. It seems to be entirely based on the inferred API cost the subscription usage could give you, but…
They are subsidized by the huge losses incurred by the AI companies.
Re: Why current LLM costs are not sustainable
#136A few thoughts: 1. Chat, being 3 yr old, is a fairly mature and solved problem today. Top companies aren't even talking about it anymore! Gemma 31B does it amazingly well (for $0.4/1M token output). Practically every near-SoTA and SoTA model does simple "chat-like" QA amazingly well -- summarization, basic question answering, single- or few-step search. 2. Tasks -- or knowledge work on a computer -- are the new front…
Re: Why current LLM costs are not sustainable
#137Earlier quoted context omitted.
Only if those losses are coming from subscriptions, instead of capex and training, which is not at all clear.
I don't understand this argument. How does it make the subscription any less subsidised if the losses are only because developing the product is just so darn expensive? Feels like arguing that it's not clear if Bugatti's losses came from selling the Veyron instead of designing and developing the Veyron.
Re: Why current LLM costs are not sustainable
#1388xB200[1] costs around 250k DIY and 450k from an enterprise builder so that will be our cost factor, these consume around 7.8kw at 100% load with median load of around 7kw (optimistic) which means that a 240kwh solar installation would be enough to supply it (72kwh buffer for bad weeks / winter) and that will set you back around $240k: this includes battery storage, installation and inverters, diy cost would be lower at around $160k.
This puts the cost of the entire system anywhere from $410k to $690k. This does not take in any property tax or land ownership into account since honestly it varies too much. The solar is simply used to provide a fixed cost basis for powering hardware instead of monthly recurring payments. Financing a 5 year loan would cost anywhere from $8,313 to $15,700.
Now let's do the math for glm-5.2[3], a fully optimized theoretical build can do around 1200tok/s which means that's around 13-14 streams of ~90tok/s on average, pushing batching further and limiting context size to around ~300k with ~150k median) you can achieve up to 37 streams at around ~40tok/s pushing performance envelope to 1400tok/s. This means you are able to generate 2.5B to 2.9B tokens in 4 weeks.
Which means putting the numbers together you can serve 1m tokens at $2.86 to $3.32 per million output tokens all else being equal. Considering that glm-5.2 is approaching opus level intelligence it's pretty safe to say that same applies for frontier labs. Input/cache write/cache reads are very difficult to price, so this assumes you're providing input / cache for free[4]. As a very heavy user I generate around 2M to 5M output tokens a day which would put me at $5.72 to $6.6 of cost per day totalling $200 a month[2].
What I also don't mention is that frontier labs have BY FAR the lowest cost per token out of any provider out there due to the amount of money they also invest into efficiency gains. This was proven by the fact that anthropic saw a huge exodus of openai users put strain on their systems and with efficiency optimizations alone they managed to mitigate a bulk of capacity issues, altho they did run into limits and had to begin spreading out the duck curve, but I have zero doubts they're getting percentage points of improvements month to month.
[1]: H300's are unobtanium unless you're building rack-rooms, H200's are not that cost effective and saturate too fast while having poorer efficiency, only capable of running flash tier models.
[2]: Okay, I didn't expect to arrive at the $200, this is kind of entertaining.
[3]: fp8, z.ai serves fp8 according to openrouter.
[4]: Assuming you want to charge for input / cache, cache reads make up roughly 30% of the cost, output 20% 50% input so to price it out it would be roughly $.3 for 1m input, $.015 for cache reads and $1.5 for output. Judging by https://openrouter.ai/z-ai/glm-5.2#pricing, appears that my math checks out.
Re: Why current LLM costs are not sustainable
#139The OpenAIs and Anthropics are going to get eaten by open source, I don't see prices going up, prices are going to crater. The models are going to be more and more commoditized.
Re: Why current LLM costs are not sustainable
#140Earlier quoted context omitted.
Opus 4.8 High effort seems adequate for me currently, at API pricing, with a $200/month budget. This is at work where I don't work on greenfield or parallelize feature development. I cannot see the agent burning through $50 for one moderately sized TypeScript cleanup in my setup. This sounds like something that can be improved on OP's side. There have been rumors about a potential Sonnet 5 model release in the near f…
> I cannot see the agent burning through $50 for one moderately sized TypeScript cleanup in my setup. Here's my usage, from the ccusage tool (slightly shortened for readability): ┌──────────┬───────────────┬────────────┬─────────────┬─────────────┬───────────────┬────────────────┬────────────────┬─────────────┐ │ Month │ Agent │ Models │ Input │ Output │ Cache Create │ Cache Read │ Total Tokens │ Cost (USD) │ ├──────…
I do not know whether that is typical, or indicative of conversations with too many turns.
Not that I would worry about this on a subscription plan, but at work where we are billed at API rates, I try to move to new conversations as often as possible.