Live data from Hacker News

Are the costs of AI agents also rising exponentially? (2025)

tobyord.com

21–30 of 155 posts

Re: Are the costs of AI agents also rising exponentially? (2025)

#21

Earlier quoted context omitted.

Selling inference is not fundamentally different from selling compute - you amortize the lifetime cost of owning and operating the GPUs and then turn that into a per-token price. The risk of loss would be if there is low demand (and thus your facilities run underutilized), but I doubt inference providers are suffering from this. Where the long-term payoff still seems speculative, is for companies doing training rathe…

There’s a lot of debate over what the useful lifespan of the hardware is though. A number that seems very vibes based determines if these datacenters are a good investment or disastrous.

I specifically remember this debate coming up when the H100 was the only player on the table and AMD came out with a card that was almost as fast in at least benchmarks but like half the cost. I haven't seen a follow up with real world use though and as a home labber I know that in the last three weeks the support for AMD stuff at least has gotten impressively useful covering even cuda if you enjoy pain and suffering.

What I'm curious about are what about the other stuff out there such as the ARM and tensor chips.

Re: Are the costs of AI agents also rising exponentially? (2025)

#22
Yet again: Transformers are fundamentally quadratic.

If they can do a task that takes 1 unit of computation for 1 dollar they will cost 100 dollars for a 10 unit task and 10,000 for a 100 unit task.

Project costs from Claude Code bear this out in the real world.

Re: Are the costs of AI agents also rising exponentially? (2025)

#23

Are any inference providers currently making profit (on inference, I know google makes money)?

All of them. It's simply impossible to sell tokens by usage at a loss now. You'll be arbitraged to death in a few days. It only makes sense to subsidize cost if you're selling a subscription.

Re: Are the costs of AI agents also rising exponentially? (2025)

#24

Interesting read. I don't know if I quite buy the evidence, but it's definitely enough to warrant further investigation. It also matches up with my personal experience, which is that tools like Claude Code are burning through more and more tokens as we push them to do bigger and bigger work. But we all know the frontier model companies are burning through money in an unsustainable race to get you and your company hoo…

> But we all know the frontier model companies are burning through money in an unsustainable race to get you and your company hooked on their tools.

Do we? Because elsewhere in the thread there's people claiming they are profitable in API billing and might be at least close to break even on subscription, given that many people don't use all of their allowance.

Re: Are the costs of AI agents also rising exponentially? (2025)

#26

Earlier quoted context omitted.

> 128GB is all you need. My guy, look around. They are coming for personal compute. Where are you going to get these 128GBs? Aquaman? [0] The ones who make RAM are inexplicably attaching their fate to the future being all LLMs only everywhere. [0] https://www.youtube.com/watch?v=0-w-pdqwiBw

Cloud can’t make money off of you and pay more than you for the hardware at the same time.

Cloud can pay more for RAM until all the RAM producers withdraw from the consumer market, then prices will go back down.

End users will still get access to RAM. The cloud terminal they purchase from Apple, Google, Samsung, or HP will have all the RAM it will ever need directly soldered onto it.

Re: Are the costs of AI agents also rising exponentially? (2025)

#27
post #26

Earlier quoted context omitted.

Cloud can’t make money off of you and pay more than you for the hardware at the same time.

Cloud can pay more for RAM until all the RAM producers withdraw from the consumer market, then prices will go back down. End users will still get access to RAM. The cloud terminal they purchase from Apple, Google, Samsung, or HP will have all the RAM it will ever need directly soldered onto it.

Doesn’t Apple place RAM directly into the SoC package? We aren’t even talking about soldering it to mother boards anymore, it is coming in with the CPU like it would as a GPU.

Re: Are the costs of AI agents also rising exponentially? (2025)

#28

Earlier quoted context omitted.

> 128GB is all you need. My guy, look around. They are coming for personal compute. Where are you going to get these 128GBs? Aquaman? [0] The ones who make RAM are inexplicably attaching their fate to the future being all LLMs only everywhere. [0] https://www.youtube.com/watch?v=0-w-pdqwiBw

Cloud can’t make money off of you and pay more than you for the hardware at the same time.

Batch inference is much more efficient. Using the hardware round the clock is much more efficient. Cloud can absolutely pay more for hardware and still make money off you.

Re: Are the costs of AI agents also rising exponentially? (2025)

#29
post #26

Earlier quoted context omitted.

Cloud can’t make money off of you and pay more than you for the hardware at the same time.

Cloud can pay more for RAM until all the RAM producers withdraw from the consumer market, then prices will go back down. End users will still get access to RAM. The cloud terminal they purchase from Apple, Google, Samsung, or HP will have all the RAM it will ever need directly soldered onto it.

I was really fucking hoping we weren't at the part where "cloud terminals" doesn't seem farfetched and paranoid and yet here we are. Jesus Christ.

Re: Are the costs of AI agents also rising exponentially? (2025)

#30

Earlier quoted context omitted.

Doubtful, local models are the competitive future that will keep prices down. 128GB is all you need. A few more generations of hardware and open models will find people pretty happy doing whatever they need to on their laptop locally with big SOTA models left for special purposes. There will be a pretty big bubble burst when there aren't enough customers for $1000/month per seat needed to sustain the enormous datacen…

Weird how you're leaving stuff like Strix Halo out. Also weird you think 128gb is the future with all of the research done to reduce that to something around 12GB being a target with all of these papers out now. I assume we'll end up with less general purpose models and more specific small ones swapped out for whatever work you are asking to do.

Strix Halo hasn‘t got nearly enough bandwidth, its just 256bit.
Post reply on HN