Live data from Hacker News

How the AI Bubble Bursts

martinvol.pe

381–390 of 557 posts

Re: How the AI Bubble Bursts

#381
post #281

Earlier quoted context omitted.

LLMs haven't remotely begun to be integrated into the lives of the typical person. Not even close. The typical person is using LLMs not at all as it pertains to their daily life tasks. They're using them almost entirely for limited discussion matters (eg having a discussion with GPT about a medical issue, or a work related matter). This is the first or second inning in the LLM rollout. It'll take 15-20 more years for…

>The typical person is using LLMs not at all as it pertains to their daily life tasks. This doesnt track at all with my experience. Everybody is using it everywhere. Moreover people are using them for daily life tasks even when it is not an appropriate use of LLMs - e.g. getting medical advice as you referred to or writing emails which are clearly pissing off their coworkers. In this respect I see it as akin to radiu…

>Everybody is using it everywhere.

No one in our Auto shop is using AI. One of the new diagnostic tools was demo'd with AI, and none of us were having it. It's about as accurate as Googling your symptoms.

My mother had an AI powered lung scan that came back with Stage 4 Cancer. The Oncologist got called in (for a fee!) to tell us it was just early stage COPD.

Re: How the AI Bubble Bursts

#382

Earlier quoted context omitted.

here you go: https://x.com/wccftech/status/2037921057097892018 Ram prices are dropping

Every response to the original post calls it out as being factually incorrect...

Nope https://x.com/rdd147/status/2038383801307639960

Re: How the AI Bubble Bursts

#383

Earlier quoted context omitted.

> The cost to serve tokens is absolutely profitable today and that’s been true for at least a year. > For the data center build outs, demand for tokens is still exceeding supply. Can you provide any numbers for this?

I can get Kimi K2.5 inference on openrouter for about $0.5/MTok input + $2.5/MTok output, from six providers that have no moat besides efficiently selling GPU time. We can assume they are doing so at a profit (they have no incentive to do this at a loss), giving us those numbers as the cost to serve a 1T-a32b model at scale. Now we don't know the true size of any of the proprietary models, but my educated guess is th…

The problem I have with this analysis is it's missing the multi-dimensional aspect of "is this profitable".

It's fair to say that if all these operators are competing for tokens, that the OpenRouter token operator (not sure the exact phrase but the people running the models) are accounting for some level of margin.

However, how many of these are running their own data centers and GPUs?

If they are running their own infrastructure, then it's not a simple equation of if each specific token set is profitable, since it needs to account for the cost of running the data center. It could be that they believe that it is profitable in the long term by utilizing the long tail of asset depreciation, but that isn't guaranteed.

IF they aren't running their own infrastructure, then it's much easier to claim that it's profitable and has a margin (outside of running their servers to manage the rented infrastructure).

HOWEVER, a lot of data centers have some pretty crazy low prices for GPUs that may be vying for user base and revenue over profitability. In these cases, if data center growth starts slowing due to slower buildout then it's very likely GPU prices go up and inference stops becoming profitable for the open router owners.

So long term it's not clear how profitable even these open models are.

OpenAI and Anthropic definitely fall into the latter category too. Their infrastructure requirements are much higher than the open models, and they are being given huge discounts so Microsoft/Amazon/Google can all claim revenue (since they have profitability coming from other parts). It's not clear if OpenAI and Anthropic models would be profitable at inference if they were paying rates that cloud hosts would make a profit from.

There's just way too many dimensions to this scenario to flat out state that open router proves inference is profitable at scale.

Re: How the AI Bubble Bursts

#384

Earlier quoted context omitted.

But this is exactly the problem - we have to take it on faith that inference is profitable because nobody actually knows. It’s hard to even define what that would mean, and while I am suspicious of claims that frontier lab CEOs are just out-and-out liars or bad people, defining and calculating the real cost of inference would be time- and labor-intensive in its own right and there is no strong incentive to do it othe…

A lot of people know. A lot of insiders have been saying tokens are profitable. Is there a conspiracy theory for everyone to lie? Including OpenAI, Anthropic CEOs, employees, Cursor management, inference providers of Chinese models?

Profitable on what basis? They generate more revenue than the cost of electricity? Does that factor in the cost to service the massive, multi-layer cake of debt that was necessary to even begin to serve inference in the first place - not from a training perspective but from a hardware and facilities perspective?

Re: How the AI Bubble Bursts

#385
I wonder if AI labs could be bailed out - like banks.

See, they kind of became a national asset and letting it go down, will leave USA watching China taking the lead for a very long time ahead. It just can't happen - right? So we'll just all fund it in taxes.

Re: How the AI Bubble Bursts

#386
post #94

Earlier quoted context omitted.

Can you give a few penciled numbers?

You can rent a H100 GPU for $4/hour. [1] 300k tokens for that hour. OpenAI charges $6. Those are pessimistic assumptions. [1] https://lambda.ai/instances

Can you keep that GPU 100% saturated at least 16 hours per day every day of the week?

If not, you aren't breaking even.

Re: How the AI Bubble Bursts

#387
post #316

Earlier quoted context omitted.

It is quite hard to imagine how the demand is saturated now. I think any company that uses a sliver of AI will happily increase their token consumption 100x if it's free.

Executive FOMO disease is being exploited by the model providers to push for maximal token usage even when it is pointless. This includes encouraging people to set up elaborate multi model set ups (e.g. "gas town") for coding that do not meaningfully improve productivity but which certainly do cause token usage to explode. It also includes encouraging execs to use token consumption as a proxy for productivity - almos…

nobody know how to measure software productivity + ai is supposed to mean productivity goes up = more ai means more productivity

As best as I can tell, that's the thinking. It's one number, it's very easy to find and manage, and there is a belief that it directly measures productivity.

I disagree that it does; seems to me the throughput of useful features is a better measure, but I'm not in the drivers seat on this one

Re: How the AI Bubble Bursts

#388

Earlier quoted context omitted.

But unlike the dot com boom, demand for tokens has not let up and there is increasing demand. I don’t know where it falls, certainly companies don’t get or right and they either over or under build. With the current demand rate changes it’s hard to understand why you would stop building today.

Demands for tokens exists yes. On one side you have huge demand for the infinitely subsidized tokens so that people can post a "unique" illustration when posting on social media, along with the text itself even. On the other end we have professionals happy to pay a subscription for heavier use, to build something in the hope to sell it. I figured I don't believe in value when my dad explained to me his mate fired his…

[dead]

Re: How the AI Bubble Bursts

#389

Earlier quoted context omitted.

That’s the wrong analogy. Model training is more like the setup costs of developing the menu and training staff. What’s driving the costs is important when talking about financial sustainability. If it’s mostly coming from optional R&D investments instead of the direct costs of producing the food then you can simply not exercise the option and be profitable. If it’s more coming as a variable cost that scales with eac…

These companies are continually training new models. This is not a long term amortized cost, it's actual COGS. Yeah, sure you can ignore the cost of purchasing the building for the restaurant for most profitability calculations, but if every year or two you were tearing down your old building and building a new one, you better believe that has to be in your profitability calculations.

I think the relevant question to define the counterfactual is what would happen if they stopped training.

If you can simply not remodel your restaurant and keep making money, then yeah, it makes sense to call it profitable.

Re: How the AI Bubble Bursts

#390
post #227

Earlier quoted context omitted.

I don't get it tbh. What market participants were speculating here? There aren't futures markets in RAM as far I know, though I certainly don't know much. And the supply constraints appear to have been pretty real (though maybe not immediate) if eg. Valve was begging publicly for RAM consignments. Were there pure-play speculators filling warehouses with DDR5?

>There aren't futures markets in RAM as far I know sure there is. not formally, but if you hold a contract for x units of future production, you can sell that contract to somebody else who wants those units more than you do.

That’s a forward contract yeah. They def do exist.

Futures are standardised forward contracts traded on exchanges

Post reply on HN