Live data from Hacker News

Why current LLM costs are not sustainable

aditya.patadia.org

121–130 of 216 posts

Re: Why current LLM costs are not sustainable

#121

Earlier quoted context omitted.

>3. We're massively overusing SOTA models. As long as you're on a subsidized subscription, you can use Claude Opus 4.8 high to write blog article meta descriptions. If you paid by token, you wouldn't do that. This idea that the subscriptions are subsidized is repeated over and over, but I've never seen any proof of this. It seems to be entirely based on the inferred API cost the subscription usage could give you, but…

> This idea that the subscriptions are subsidized is repeated over and over, but I've never seen any proof of this. It seems to be entirely based on the inferred API cost the subscription usage could give you, but there are a lot of assumptions needed for that to follow. My claude code environment shows me cost per token used in that session, according to API costs. It regularly exceeds $200. I pay $200 a month for m…

But they also determine the token prices. What you describe could also be true if they take a 5x profit margin on api tokens and 2x margin on subscriptions.

Re: Why current LLM costs are not sustainable

#122
post #51
post #3

I have already seen a number of people doing the math on what it would take for hardware to self host a Q8XL quantization of GLM5.2 shared between N numbers of people. There's additional advantages that everything you query, all of your context cache and everything it outputs stays private and can't be arbitrarily turned off by external interference. Personally I think it would be a fairly good bet that something wit…

Back in the earlier days of the internet, when "dedicated servers" were a competitive advantage, hobbyists and small dev shops definitely shared dedicated hardware. So you could see small LLM co-operatives working out, yeah. But my thinking is that this four-to-five-year scenario just won't come to fruition, because the whole concept of needing to run these massive, massive models will slightly more likely be rendere…

the biggest problem is most ai will be local by 2030. Every future device you buy will have AI compute on it somewhere, built in, like it has on-device floating point.

On top of this, people are constantly coming up with better ways of running models on less special hardware and "good enough" models are now existing for most tasks.

So where does that leave the frontier labs? Drug discovery? Maybe some hard math problems? I mean it's not that big actually...

We're in a brief window where this is profitable, like batch computing was in the 70s. However, once your own device can do it, you're going to start migrating.

Re: Why current LLM costs are not sustainable

#123

Earlier quoted context omitted.

>3. We're massively overusing SOTA models. As long as you're on a subsidized subscription, you can use Claude Opus 4.8 high to write blog article meta descriptions. If you paid by token, you wouldn't do that. This idea that the subscriptions are subsidized is repeated over and over, but I've never seen any proof of this. It seems to be entirely based on the inferred API cost the subscription usage could give you, but…

> This idea that the subscriptions are subsidized is repeated over and over, but I've never seen any proof of this. It seems to be entirely based on the inferred API cost the subscription usage could give you, but there are a lot of assumptions needed for that to follow. My claude code environment shows me cost per token used in that session, according to API costs. It regularly exceeds $200. I pay $200 a month for m…

That's what they want to charge you. Not the actual cost. The actual cost is a gpu that's probably already paid off and about $2 of electricity

Re: Why current LLM costs are not sustainable

#124
post #15

There is a wave of users switching over to DeepSeek Flash. There are Reddit threads of users sharing billion token spend for $20. If all of global spend on Anthropic/OpenAI/Gemini APIs just switches over to DeepSeek then easily we can decrease total AI spend by 10x

Probably won't be too long before the government decides to block deepseek's website based on "security" concerns.

Deepseek's models are open-weight and hosted all over the world, how would blocking deepseek's web sight do anything to stop its model's use?

Re: Why current LLM costs are not sustainable

#125
post #24

Earlier quoted context omitted.

I think this comes from the idea that serving these tokens without paying for training is already expensive, e.g. https://news.ycombinator.com/item?id=46613887 self-hosted solution might give you only 10-100x more affordable solution at cost . So, given the SOTA providers with even larger models also need to continously be using considerable resources for training their next models, to fund future data centers, and m…

Except there are plenty of inference providers worldwide (including the US) that serve open-weight models that are not subsidized, and are reasonable in cost. Or is your claim that those are all running at a loss?

So they do not train models, and in addition their models are expected to be smaller than SOTA models, although we cannot know for sure by how much.

So what's the price difference, 3000x?

Re: Why current LLM costs are not sustainable

#126
post #9

There is a wave of users switching over to DeepSeek Flash. There are Reddit threads of users sharing billion token spend for $20. If all of global spend on Anthropic/OpenAI/Gemini APIs just switches over to DeepSeek then easily we can decrease total AI spend by 10x

I am not sure if that is wise. It’s a hostile superpower after all

I'm in Europe. The only superpower that's been hostile to me, very directly - was US, when they asked a company I was relying on to limit model access based on nationality.

China has (so far), never done that to me.

Re: Why current LLM costs are not sustainable

#127
post #118

It's weird to see people claiming that model capabilities are plateauing. It wasn't until late last year that we even had strong coding models. Imagine if, less than a year after the first iPhone launched, people claimed that smartphone capabilities were "plateauing" because Apple hadn't yet launched a new phone. And it seems the issue is less than "models aren't getting better" than, "models are good enough to handl…

People claim what they see. I see no improvement since opus 4.6, quite the opossite.

Which is not even 5 months old.

Re: Why current LLM costs are not sustainable

#128
post #6

> To give an example, just doing Typescript type fixes with this model across 50 files cost me $54 this afternoon. If you can use a subscription with any of the SOTA models, do that. Instead of around 4k EUR in token costs, my Opus usage costs me 108 EUR (with taxes) per month with their Max 5x plan. It's the same with OpenAI, those are heavily subsidized. It doesn't make sense to pay per-token, unless you must. > Wh…

I think it will end up a game of haves and have nots - governments and companies with spare cash will go large and capitalise on the benefits of ai, everyone else will be left behind

Re: Why current LLM costs are not sustainable

#129
post #3

I have already seen a number of people doing the math on what it would take for hardware to self host a Q8XL quantization of GLM5.2 shared between N numbers of people. There's additional advantages that everything you query, all of your context cache and everything it outputs stays private and can't be arbitrarily turned off by external interference. Personally I think it would be a fairly good bet that something wit…

the same argument was made 2 years ago: "in 2 years we'll be able to run GPT-4 level models on an expensive laptop, most people will be using this instead of the fancy cloud models".

we are there, Gemma4/Qwen3.6 are GPT-4 level models runnable on a fancy laptop.

but expectations shifted, nobody wants a GPT-4 level model anymore

Re: Why current LLM costs are not sustainable

#130
As information flows abundantly, and as information processing flows more abundantly, where will the bottlenecks in the system emerge? It surely won't be in design and production. It probably won't be in chips, infrastructure, and energy (already commodities, increasingly competitive). So... what's the bottleneck? Political will? Human discernment/taste? Raw materials?
Post reply on HN