Live data from Hacker News

AI's Affordability Crisis

blog.dshr.org

211–220 of 436 posts

Re: AI's Affordability Crisis

#211

This is basically bunk because AI costs have gone down by 50x or more (api costs) since 3 years.

I'm no economist but if true don't you have the opposite problem? How do you get people to need X many tokens per day such that you can sell enough to make money? Wouldn't you need an absence of competition for that to be ok?

the demand for intelligence is infinite. you sound like someone in 1960 wondering what the hell we would even do with the functionally infinite cpu cycles we have available to us now.

Re: AI's Affordability Crisis

#212

The article fails to mention DeepSeek, Alibaba, Qwen, Xiaomi, MiMo, z.ai, or GLM. It's hard to take such an article seriously that doesn't do this. (Our monthly total spend is around $180 with a team of 6, about half technical; our biggest line items are for American models or subscriptions which we probably will be planning to get rid of.) And then remarks like this: Anthropic, OpenAI and Microsoft have all now tran…

> Our monthly total spend is around $180 with a team of 6, about half technical; our biggest line items are for American models or subscriptions which we probably will be planning to get rid of.) Please tell more :). Do you pay per token from bedrock / openrouter / somewhere else? How many tokens you use over the month, and how many for each task? Which harnesses?

Pay for DeepSeek directly. One developer insists on having his own account and in theory expenses it, but he forgets to turn in $10 expense reports. (Total spend in last two months = about $45.)

Pay for OpenAI Pro directly, but I’m the only guy that uses Codex. $100 a month. My nontechnical partner likes to talk to ChatGPT 5.5 Pro for image related tasks (think generating interior decorating pics).

The nontechnical staff use a Gemini account on a Google family AI Pro sub. I use Antigravity when working on Android or Google Cloud API codebases.

Everyone gets OpenCode Go. The cost is trivial. $10 a month per person.

Pay for MiMo directly. We use it during Chinese off peak hours though. Total spend so far $25 in last month.

We run a few Qwen models locally and pretty much have them pegged all day. RTX 5090 on a PC and a Mac Studio.

There’s also Grok which is used for Imagine for artistic / graphic design related work. I also use the subscription for a vision model in my oh-my-pi harness.

We’re having discussions about how to pull in GLM-5.2 cost effectively. We compete with third world development shops so we can’t really pass on inference costs, but we can benefit from getting jobs done for customers faster. But ⅔ of our work is either internal or open source projects we can’t bill for.

Re: AI's Affordability Crisis

#213

Affordability is not the current goal. Vendor lock-in is the current goal. Consumer prices are a drop in the bucket comparatively.

How can you lock in when the harnesses are basically thin clients around the APIs and you can replicate them using agents in a short period of time? I haven't seen a compelling thesis yet for how you achieve vendor lock in for LLMs. Claude Code is a bit sticky, but if we're being honest its just because Codex doesn't have all the same features yet.

Re: AI's Affordability Crisis

#214
post #127

Earlier quoted context omitted.

The big thing is, the western world has moved so much of the manufacturing to China and think a lot of people will not forgive Samsung and others, so I can see China owning a good portion of the supply chain.

> The big thing is, the western world has moved so much of the manufacturing to China I built my career on Solaris and it got rugpulled by Linux. That wasn’t because of software, it was because of hardware. Linux’s cost advantage existed because Sun hardware had huge margins, because their software was basically free. AI will probably be a repeat of this. Whoever can come up with the hardware solution that minimizes…

but not all tokens are equal and vertical integration is the name of the game. Solaris did not lose to Linux, it lost to the LAMP stack on commodity x86 hardware. without the "AMP" part, Linux would've been dead in the water.

Re: AI's Affordability Crisis

#215

Earlier quoted context omitted.

I saw some commentary that their free cash flow is misleading because it doesn't subtract the stock compensation they are paying to attract / keep top AI talent. Their point was also that deciphering financial statements is hard

Why would it? Stock compensation doesn't affect cash flow, it just dilutes the shareholders.

Except that's the thing, they do stock buybacks so they do not dilute existing shareholders or lower stock prices.

This is the video I watched that explained the shenanigans (from the guests' perspective, not illegal, obfuscated)

https://www.youtube.com/watch?v=YrJzjC4kKCY

Re: AI's Affordability Crisis

#216

Earlier quoted context omitted.

You won't need a frontier size model for most tasks before long. Qwen 3.6 (small) punches way above its weight. I run it at home @8bit on an OEM Spark

And corporations could run DeepSeek models on cloud hardware.

You can run most open models on cloud hardware. Google Cloud gives you a click to deploy, but then you have saturation / ROI considerations, versus Google serving them up multi-tenant, per-token.

Re: AI's Affordability Crisis

#217
post #191

Earlier quoted context omitted.

I think that for coding we're past the plateau issue. The frontier models of today are good enough and very valuable. The expensiveness in running them will eventually be solved by cheaper faster hardware. I do hope that a day will come where you can buy the nvidia spark thingy for 5k that can run the equivalent of Opus 4.6 or 4.5 locally and that would be a massive thing.

> The expensiveness in running them will eventually be solved by cheaper faster hardware. How? * Moores Law is almost over. The 5090 improves over the 4090 mostly because of quant improvements. * even if the hardware improves, there’s a huge incentive to slow roll the next generation. Nobody wants to end up like Sun Microsystems. Sun’s used hardware was faster than its new hardware, once you considered price. Sun end…

GPUs are not really the ideal architecture for running neural networks; they are heavily bottlenecked by memory bandwidth and struggle to keep all their tensor cores supplied with data.

There is significant room to make more specialized neural network accelerators with new compute-in-memory architectures.

If the brain can run 86 billion neurons on 30W it must be possible.

Re: AI's Affordability Crisis

#218
It's not an affordability crisis, it's a financial crisis. The models get cheaper super fast. By this time next year Fable 5 will cost less than Sonnet does today. That's not the problem. The problem is that many companies are going to realize that they don't get any ROI from AI. Generating code faster != more profit. Most of the Fortune 500 will likely realize this and then the token budgets will come crashing down. Most of their ideas are _bad_ ideas. Implementing bad ideas faster, won't lead to more profit.

Sure, you can use AI to potentially replace software engineers, but the F500 are also terrified of not having accountability or making mistakes. They won't be firing any engineers. In that scenario, there's just no room for AI usage. If you have to be responsible for all the code, then... AI has to either manage it completely autonomously (which even Fable can't) or... humans have to be in the loop which means they still have to understand the code. The best way to understand the code is to write the code yourself. So there's no productivity gain to be had.

I'm pro-AI, but I think we're due for a big crash next year.

Re: AI's Affordability Crisis

#219

Earlier quoted context omitted.

The US govt is going to ban foreign models and foreign providers, and frontier labs are still cooked, because US companies will RLwash Chinese models to try and get in on the captive market. The frontier labs have already lost the war for coding, their next play is custom models for specific domains... Anthropic Galen for biomedical research, Anthropic Locke for legal analysis, etc, and you won't see _ANY_ intermedia…

> The US govt is going to ban foreign models The people have a right to make and use whatever models they want, protected by the constitution. At a minimum, the models are described in research papers that are unquestionably protected speech. Skilled devs turn those into programs, also protected speech.

> protected by the constitution

I don't see how.

Re: AI's Affordability Crisis

#220

Deepseek is 90% cheaper, and nearly as good for coding tasks as claude/codex, and as good given the right plan. The only moat OpenAI and Anthropic have is regulation. If the Chinese really eant to hammer us, they could realse the full training data and pipeline.

Even without doing that the Chinese are already going to impact our labs presence everywhere else in the world. With Fable getting pulled, any model coming out of the US is now unreliable and untrusted. No one in any other country would in their right mind choose OpenAI or Anthropic for anything. The big push for regulation and export controls is only going to ensure OpenAI & Anthropic are more like the automakers. O…

I have to push back on this: China's cheap EVs and power prices are due to industrial policy on an epic scale which goes directly against the whole free market thing. I personally think industrial policy is a good thing, but you cannot have it both ways and not expect workers to get unhappy and vote against your interests when they have no more jobs.
Post reply on HN