Earlier quoted context omitted.
>3. We're massively overusing SOTA models. As long as you're on a subsidized subscription, you can use Claude Opus 4.8 high to write blog article meta descriptions. If you paid by token, you wouldn't do that. This idea that the subscriptions are subsidized is repeated over and over, but I've never seen any proof of this. It seems to be entirely based on the inferred API cost the subscription usage could give you, but…
> This idea that the subscriptions are subsidized is repeated over and over, but I've never seen any proof of this. It seems to be entirely based on the inferred API cost the subscription usage could give you, but there are a lot of assumptions needed for that to follow. My claude code environment shows me cost per token used in that session, according to API costs. It regularly exceeds $200. I pay $200 a month for m…
Why current LLM costs are not sustainable
121–130 of 216 posts
Re: Why current LLM costs are not sustainable
#122I have already seen a number of people doing the math on what it would take for hardware to self host a Q8XL quantization of GLM5.2 shared between N numbers of people. There's additional advantages that everything you query, all of your context cache and everything it outputs stays private and can't be arbitrarily turned off by external interference. Personally I think it would be a fairly good bet that something wit…
Back in the earlier days of the internet, when "dedicated servers" were a competitive advantage, hobbyists and small dev shops definitely shared dedicated hardware. So you could see small LLM co-operatives working out, yeah. But my thinking is that this four-to-five-year scenario just won't come to fruition, because the whole concept of needing to run these massive, massive models will slightly more likely be rendere…
On top of this, people are constantly coming up with better ways of running models on less special hardware and "good enough" models are now existing for most tasks.
So where does that leave the frontier labs? Drug discovery? Maybe some hard math problems? I mean it's not that big actually...
We're in a brief window where this is profitable, like batch computing was in the 70s. However, once your own device can do it, you're going to start migrating.
Re: Why current LLM costs are not sustainable
#123Earlier quoted context omitted.
>3. We're massively overusing SOTA models. As long as you're on a subsidized subscription, you can use Claude Opus 4.8 high to write blog article meta descriptions. If you paid by token, you wouldn't do that. This idea that the subscriptions are subsidized is repeated over and over, but I've never seen any proof of this. It seems to be entirely based on the inferred API cost the subscription usage could give you, but…
> This idea that the subscriptions are subsidized is repeated over and over, but I've never seen any proof of this. It seems to be entirely based on the inferred API cost the subscription usage could give you, but there are a lot of assumptions needed for that to follow. My claude code environment shows me cost per token used in that session, according to API costs. It regularly exceeds $200. I pay $200 a month for m…
Re: Why current LLM costs are not sustainable
#124There is a wave of users switching over to DeepSeek Flash. There are Reddit threads of users sharing billion token spend for $20. If all of global spend on Anthropic/OpenAI/Gemini APIs just switches over to DeepSeek then easily we can decrease total AI spend by 10x
Probably won't be too long before the government decides to block deepseek's website based on "security" concerns.
Re: Why current LLM costs are not sustainable
#125Earlier quoted context omitted.
I think this comes from the idea that serving these tokens without paying for training is already expensive, e.g. https://news.ycombinator.com/item?id=46613887 self-hosted solution might give you only 10-100x more affordable solution at cost . So, given the SOTA providers with even larger models also need to continously be using considerable resources for training their next models, to fund future data centers, and m…
Except there are plenty of inference providers worldwide (including the US) that serve open-weight models that are not subsidized, and are reasonable in cost. Or is your claim that those are all running at a loss?
So what's the price difference, 3000x?
Re: Why current LLM costs are not sustainable
#126There is a wave of users switching over to DeepSeek Flash. There are Reddit threads of users sharing billion token spend for $20. If all of global spend on Anthropic/OpenAI/Gemini APIs just switches over to DeepSeek then easily we can decrease total AI spend by 10x
I am not sure if that is wise. It’s a hostile superpower after all
China has (so far), never done that to me.
Re: Why current LLM costs are not sustainable
#127It's weird to see people claiming that model capabilities are plateauing. It wasn't until late last year that we even had strong coding models. Imagine if, less than a year after the first iPhone launched, people claimed that smartphone capabilities were "plateauing" because Apple hadn't yet launched a new phone. And it seems the issue is less than "models aren't getting better" than, "models are good enough to handl…
People claim what they see. I see no improvement since opus 4.6, quite the opossite.
Re: Why current LLM costs are not sustainable
#128> To give an example, just doing Typescript type fixes with this model across 50 files cost me $54 this afternoon. If you can use a subscription with any of the SOTA models, do that. Instead of around 4k EUR in token costs, my Opus usage costs me 108 EUR (with taxes) per month with their Max 5x plan. It's the same with OpenAI, those are heavily subsidized. It doesn't make sense to pay per-token, unless you must. > Wh…
Re: Why current LLM costs are not sustainable
#129I have already seen a number of people doing the math on what it would take for hardware to self host a Q8XL quantization of GLM5.2 shared between N numbers of people. There's additional advantages that everything you query, all of your context cache and everything it outputs stays private and can't be arbitrarily turned off by external interference. Personally I think it would be a fairly good bet that something wit…
we are there, Gemma4/Qwen3.6 are GPT-4 level models runnable on a fancy laptop.
but expectations shifted, nobody wants a GPT-4 level model anymore