Live data from Hacker News

Are the costs of AI agents also rising exponentially? (2025)

tobyord.com

111–120 of 155 posts

Re: Are the costs of AI agents also rising exponentially? (2025)

#111

Earlier quoted context omitted.

With some research, that chip appears like it would cost about $300-$400 to manufacture, die only. For an 8B parameter model. Opus is estimated at 500B-2T parameters. At that scale you’re past reticle limits and need HBM and multi-die packaging, which means you’ve essentially built an inference ASIC (like Groq or Etched) rather than something categorically cheaper than GPUs. The “burned into silicon” advantage mostly…

The cutting edge, max size models will likely stay in the GPU space for a long time. But these models are not needed for most general requests. With a fine tuned 30B quantisized model you can serve a large portion of requests with around 32GB of RAM. Free users will likely only get these kinds of models. At some point we will get these models in hardware and the cost per token will be minimal.

> With a fine tuned 30B quantisized model you can serve a large portion of requests with around 32GB of RAM. Free users will likely only get these kinds of models.

These are exactly the kinds of models that you can easily run locally by repurposing existing hardware. Depending on how much you're willing to wait for the answer, running local even gives you strictly better outcomes for simple Q&A queries.

(Long-context and agentic use cases are admittedly much harder to fit under that model, since non-AI uses for the high-end hardware you'd realistically need for those are rather more limited, and they're hit by the ongoing hardware shortage.)

Re: Are the costs of AI agents also rising exponentially? (2025)

#112

The sweet spot thing is the real insight here and nobody seems to be talking about it. Frontier models get hyped for their maximum task horizon, but that's also where they're 10-30x more expensive per hour than their optimal range. You're paying a massive premium for the hardest tasks and still failing half the time. Honestly the practical takeaway is pretty boring: just break your work into smaller chunks. Not becau…

Small chunks of work start to become viable for local agentic use too. The O(N^2) dependence on context length really makes the "maximum tasks" a complete non-starter locally.

Re: Are the costs of AI agents also rising exponentially? (2025)

#113
post #78

The crazy part about this is if you compare it not to US wages but european, for instance in the UK where the median software hourly wage is somewhere around $35-40 an hour, then humans are already cheaper than the best models.

Humans are not cheaper than AI models. Let's go with $35 an hour. 24 365 = 8760 8760 $35 = $306,600 Yeah, a human working non stop will run $300k. Now you said, the "best" models. I personally reckon that 80-90% of most work don't need the best models. They need a good model, and good models are super cheap. i.e, the tiny gemma4 or qwen3.6 models will be sufficient for most of those work. AI cloud usage cost goes up…

> So say someone built an under $10k system, with perhaps dual RTX 5090. That same system will be able to easily run 20 parallel requests. The only cost is electricity. You can run it 24/7. For 1 year, that's ~$6million

I dont see how you get anywhere close to $6M of tokens out of a pair of 5090s. The class of model they could run is fairly small and extremely cheap to run via API (my math says running Gemma4-31B for 24 hours costs less than $1 on OpenRouter). Even with 20x concurrent requests you are orders of magnitude away from $6M/yr.

Re: Are the costs of AI agents also rising exponentially? (2025)

#114
post #46

My expectation: demand going up, prices will rise, supply will saturate to the point of ubiquitous "utility" status, and prices will drop, probably a bell curve shape with sine-wave undulations along the way.

> supply will saturate that depends on the ability to produce supply at a saturation rate. It did work for internet backhaul links - ala, those dark fibres. However, i reckon those fibres are easier to manufacture than silicon chips. I wonder if saturation is possible for ai capable chips.

no concerns for the hardware chain long view. saturate is AI as utility ubiquitous in everything.

Re: Are the costs of AI agents also rising exponentially? (2025)

#115
post #63
post #38

Earlier quoted context omitted.

Anthropic has 50% gross margins on their tokens. Step 1) Bubble callers will be proven wrong in 2026 if not already (no excess capacity) Step 2) Models are not profitable are proven wrong (When Anthropic files their S1) Step 3) FOMO and actual bubble (say around 2028/29)

If they had such a high margin, they wouldn't need to fuck around with token usage/pricing every three days. I have no data to support this, but I think they just about break even on API usage and take overall loss on subscriptions/free plans.

Math / Economics 101 thought experiment.

You have (limited) 100 Coke cans to sell (that you bought for say $1)

There are two large lines being formed for that. One line is offering an average $3 per bottle and another line is offering an average $2 per bottle.

Tell me which line they would throttle/starve even though they make a profit out of it.

Also, when the lines were formed you had no idea of the average price, but now you are getting a clear picture. Would you change your strategy / pricing or stick with your original "give the bottle to everyone for the same initial $1 price"

Re: Are the costs of AI agents also rising exponentially? (2025)

#116

The sweet spot thing is the real insight here and nobody seems to be talking about it. Frontier models get hyped for their maximum task horizon, but that's also where they're 10-30x more expensive per hour than their optimal range. You're paying a massive premium for the hardest tasks and still failing half the time. Honestly the practical takeaway is pretty boring: just break your work into smaller chunks. Not becau…

Model specialization is in all likelihood going to be the way forward, both for cost and quality of output. Smaller, cheaper models specialized in their task domains. Many of the current model vendors are already (attempting) to do this under the hood. Generalist models have similar problems as generalist humans. The proverbial "Jack of all trades, master of none." That said, I've made my career as a generalist :)

Anyone trying to decide which of 30 different specialized models best fits their task has already failed.

Maybe the future of the backend is specialized models but the future of what faces the user is what appears to be a generalist model. Maybe it does things itself, maybe it just knows how to route to the specialist models, but the UX of a generalist model will win.

Re: Are the costs of AI agents also rising exponentially? (2025)

#117

Earlier quoted context omitted.

No, burning models into hardware won't make them faster or reduce the cost. It will cost way more for similar performance as what you would get with a gpu. I am not telling you why, you can go figure that out on your own.

But isn't this happening here https://taalas.com/ already. They have a demo of llama running at 17000 tokens per second https://chatjimmy.ai/

You mean the person saying "I won't tell you why" might not know what they're talking about?! Say it ain't so.

Re: Are the costs of AI agents also rising exponentially? (2025)

#118
post #38

Earlier quoted context omitted.

Anthropic has 50% gross margins on their tokens. Step 1) Bubble callers will be proven wrong in 2026 if not already (no excess capacity) Step 2) Models are not profitable are proven wrong (When Anthropic files their S1) Step 3) FOMO and actual bubble (say around 2028/29)

Can we see them?

https://www.theinformation.com/articles/anthropic-lowers-pro...

I have access to that article

https://www.saastr.com/have-ai-gross-margins-really-turned-t...

Like I said, majority of people (including smart ones) are going to be surprised by the profit margins of AI labs and there will be a mad rush to buy AI stocks till it reaches bubble proportions.

2025 was merely a 1996 "Irrational Exuberance" moment. We haven't seen the late 1999 mania yet

Re: Are the costs of AI agents also rising exponentially? (2025)

#119

Once a model is stable and good enough, for example Sonnet 4.6 or GPT 5.4 (or something else in future), it can be burned into hardware like Talaas chip reducing the cost many times and increasing the speed. At some point we can rely on old model while being productive with it.

I always wondered why the equivalent of integrated mining didn't apply to LLM inference... now it turns out it does and there's a company making it fast and robust!

Re: Are the costs of AI agents also rising exponentially? (2025)

#120

Earlier quoted context omitted.

I tried to use gpt for various handy work. While it does help I don't think it can adequately substitute for hard won hands. Maybe next gen if you provide a video stream and the llm can view the exact situation. Even then though I wouldn't discount the difficulty of learning dexterity when you've been a coddled white collar worker your whole life

I wasn't suggesting white collar workers attempt blue collar work. I'm merely saying that cheap day laborers with basic experience won't have to lean on their industry mentorship model (journeyman etc) as much and can complete jobs on their own. On the cheap. Today's models are insufficient for someone with 0 hands on experience, especially when limited to text modalities. However, I don't doubt the future ones you d…

Is your opinion here grounded in experience from working in that field, or is it speculation?
Post reply on HN