Live data from Hacker News

AI's Affordability Crisis

blog.dshr.org

251–260 of 436 posts

Re: AI's Affordability Crisis

#251

Earlier quoted context omitted.

The US government isn’t supposed to be allowed to constrain speech, but they do have the power to constrain commerce, and they can ban the sale of AI services and AI-capable hardware if they choose.

> but they do have the power to constrain commerce its an interesting idea; i'd like to see someone claim buying/selling as a form of speech...

Citizens United got pretty close in the Supreme Court's 5–4 decision on January 21, 2010, ruling that corporations and unions cannot be prohibited from making independent political expenditures, citing First Amendment free speech protections.

Re: AI's Affordability Crisis

#252

Earlier quoted context omitted.

I would imagine it only gets worse in the face of good-enough open/chinese/local models too right? Microsoft adding Deepseek support already as I recall? That is - for any definition of "they are behind X months" then eventually they get to the point Claude was in January when the world freaked out, but at 1/10th the cost. A lot of firms are going to mandate that is good enough for their developers.

Yup, we are in the process of getting access to US hosted Chinese models. I've been petitioning Google and our rep, we will see but I suspect they will cave eventually. Gemini sucks and if they don't sell what their customers want, we go shopping around.

What you want is already available on OpenRouter and a million other services, but sure, you can wait 18 months for it to be on GCP.

Re: AI's Affordability Crisis

#253

Earlier quoted context omitted.

You can run most open models on cloud hardware. Google Cloud gives you a click to deploy, but then you have saturation / ROI considerations, versus Google serving them up multi-tenant, per-token.

For ROI though you can run 24/7 agentic-style workloads, constantly churning through all your source code looking for security bugs (or whatever) and you DONT pay per-token costs. A DeepSeek instance running 24/7 in a cloud provider will beat doing that with Claude which could bankrupt you with 100x more costs, even though it might find more. And DeepSeek may find enough to keep your engineering team saturated and bu…

This process works better with multiple models and not simply slinging Ai at it 24/7. It's not an ROI just because you keep the GPU busy. The signal to noise ratio is not there yet

Re: AI's Affordability Crisis

#254
post #252

Earlier quoted context omitted.

Yup, we are in the process of getting access to US hosted Chinese models. I've been petitioning Google and our rep, we will see but I suspect they will cave eventually. Gemini sucks and if they don't sell what their customers want, we go shopping around.

What you want is already available on OpenRouter and a million other services, but sure, you can wait 18 months for it to be on GCP.

> We are in the process...

OpenRouter charges an extra 5.5%, Fireworks does not, Google is separate, but I doubt it will take 18 months. They are already aware they are losing business.

OpenRouter is the wrong abstraction for enterprise, we only need one model provider, not everyone in the world. Nor do we want to have to worry about failover going to providers we don't want.

Re: AI's Affordability Crisis

#255

Earlier quoted context omitted.

> The expensiveness in running them will eventually be solved by cheaper faster hardware. How? * Moores Law is almost over. The 5090 improves over the 4090 mostly because of quant improvements. * even if the hardware improves, there’s a huge incentive to slow roll the next generation. Nobody wants to end up like Sun Microsystems. Sun’s used hardware was faster than its new hardware, once you considered price. Sun end…

GPUs are not really the ideal architecture for running neural networks; they are heavily bottlenecked by memory bandwidth and struggle to keep all their tensor cores supplied with data. There is significant room to make more specialized neural network accelerators with new compute-in-memory architectures. If the brain can run 86 billion neurons on 30W it must be possible.

Our brains run 86 billion neurons the same way a waterfall runs a fluid simulation with N quadrillion particles.

Re: AI's Affordability Crisis

#256
post #71

> Zitron's numbers don't tell us the real cost of generating tokens but, subject to the assumption that the platforms are not subsidizing the token price, that means Anthropic is subsidizing their enterprise customers by up to 40 times, and OpenAI up to 70 times Neither Anthropic nor OpenAI are subsidizing enterprise customers. Neither Anthropic nor OpenAI allow Business nor Enterprise customers access to the high va…

> it may underestimate the inference margins of Ant/OAI's API pricing. If true then why are neither Anthropic or OpenAI dropping their API pricing to gain market share when both are clearly doing all sorts of political and PR maneuvering to compete in a cutthroat market? Since they aren't dropping the API usage prices (and are in fact raising them in a lot of subtle ways) then one of these options almost has to be tr…

> If true then why are neither Anthropic or OpenAI dropping their API pricing to gain market share

Maybe because they're trying to IPO this year, and their IPO prospects will be a lot worse if their S-1s show them to be losing money on inference as opposed to making a healthy profit.

Re: AI's Affordability Crisis

#257
post #49

Earlier quoted context omitted.

I.e., the demand for programming tokens turns out to be quite elastic.

I would imagine it only gets worse in the face of good-enough open/chinese/local models too right? Microsoft adding Deepseek support already as I recall? That is - for any definition of "they are behind X months" then eventually they get to the point Claude was in January when the world freaked out, but at 1/10th the cost. A lot of firms are going to mandate that is good enough for their developers.

I'm set up to use Qwen 3.6 locally if needed. It's solid, it does what I need, it runs on my laptop and it's free.

But that's because I never got on the "run three dozen agents in a ralph loop" trend or other high-token usage methods. The way I use AI is discrete and targeted and it seems that's how it will be for everyone once the economics settle.

Re: AI's Affordability Crisis

#258

I think the biggest problem is not necessarily the cost to develop & serve the models, but how quickly user behavior changed with token based pricing. I know a lot of people at companies where the marching orders changed on a dime end of Q1/start of Q2. These are shops that were fully on the "use AI or die (because we will fire you)" train. Now there's monitoring, reporting, alerting not just on overall cost but on "…

Our company went from “AI AI AI” to “GitHub Copilot has been suspended due to exceeding the budget” with this month’s price increase.

Re: AI's Affordability Crisis

#259

Earlier quoted context omitted.

Why would it? Stock compensation doesn't affect cash flow, it just dilutes the shareholders.

Except that's the thing, they do stock buybacks so they do not dilute existing shareholders or lower stock prices. This is the video I watched that explained the shenanigans (from the guests' perspective, not illegal, obfuscated) https://www.youtube.com/watch?v=YrJzjC4kKCY

Yeah but that’s very standard and pretty much all pre-profit tech companies do something like that when/if they can

Re: AI's Affordability Crisis

#260
post #143

Lol I feel like no one has any attention span here. Tech shit is expensive in the beginning when it's new. It gets cheaper with time. This is a tech forum, don't we know this? Of course people overreact in both directions on both sides of the issue. It's a very fast technology, wait for things to settle before making grand declarations.

> Lol I feel like no one has any attention span here. Tech shit is expensive in the beginning when it's new. It gets cheaper with time. The funniest comment here. Have you seen the prices of the technical shit for the past two years? Dang, GPUs are not getting any cheaper, but more expensive with each year.

It’s a massive supply crunch. More production will come online.
Post reply on HN