Live data from Hacker News

Why current LLM costs are not sustainable

aditya.patadia.org

11–20 of 216 posts

Re: Why current LLM costs are not sustainable

#11
I am convinced that the combination of capable open weight models and specialized hardware will mean that Apple (and other hardware providers) will start shipping computers with built-in, hardwired, "LLM-on-a-chip" cards that are capable enough to meet 90% of your AI needs.

I really believe that in the near-term future we will run our LLMs in hardware, not in software. Hardwire a capable model into a device the size of a graphics card, embed it into a laptop, and you have something that uses less power, does faster inference, doesn't require additional CPU or memory, doesn't cost a monthly fee, and will probably eventually be available for under a (few) hundred bucks.

Re: Why current LLM costs are not sustainable

#12
post #9

There is a wave of users switching over to DeepSeek Flash. There are Reddit threads of users sharing billion token spend for $20. If all of global spend on Anthropic/OpenAI/Gemini APIs just switches over to DeepSeek then easily we can decrease total AI spend by 10x

I am not sure if that is wise. It’s a hostile superpower after all

Well... Open weights on premise is politically neutral.

Re: Why current LLM costs are not sustainable

#13

I think companies will fire 5-10% of people and convert them to token budget. I also believe that before any real companies are running these models locally, they will already have some kind of agentic layer. With the current frontier model lab progress, i do not see any real company which makes real money, running local models. Running local models is easy for me, for sure not that easy for any company. Your DC need…

I am calling it now. LLM hosting is the new web hosting. You will have a market of hosting providers offering you access to LLM compatible hardware (the Hetzners of the LLM world) as well as virtualised LLM access (the Heroku of the LLM world). These will compete along pricing, ownership axes while frontier labs will compete mostly on performance, integration and ease of use (think Wordpress).

That's the only way I can see frontier labs charging high enough to sustain the cash flow needed to operate as racing to the bottom is not possible for them.

It is interesting to think whether this is another "Cambrian" era like the smartphone OSes when you had Symbian, Android, iOs, Windows Mobile and so many others competing.

Re: Why current LLM costs are not sustainable

#14
The current costs do not have to be sustainable for the SOTA model providers as they grow their user base. But I really wonder about the future as the costs have to increase at some point (to be sustainable) but at the same time the competition and local models get better and better.

Re: Why current LLM costs are not sustainable

#15

There is a wave of users switching over to DeepSeek Flash. There are Reddit threads of users sharing billion token spend for $20. If all of global spend on Anthropic/OpenAI/Gemini APIs just switches over to DeepSeek then easily we can decrease total AI spend by 10x

Probably won't be too long before the government decides to block deepseek's website based on "security" concerns.

Re: Why current LLM costs are not sustainable

#18
post #9

There is a wave of users switching over to DeepSeek Flash. There are Reddit threads of users sharing billion token spend for $20. If all of global spend on Anthropic/OpenAI/Gemini APIs just switches over to DeepSeek then easily we can decrease total AI spend by 10x

I am not sure if that is wise. It’s a hostile superpower after all

Hostile? Us or them? I beg to differ who the hostile ones might be.

Re: Why current LLM costs are not sustainable

#20
post #16

Earlier quoted context omitted.

Well... Open weights on premise is politically neutral.

Try doing it at scale for a whole office. Not trivial.

You could probably do with couple of instances. People rarely use ai 24/7, so right now you can oversubscribe and still have acceptable latency and high utilization rate.
Post reply on HN