Meanwhile datacenters put out more pollution and use more electricity than all the plane rides Bill Gates took with Epstein combined, for business meetings of course.
AI subscriptions are a ticking time bomb for enterprise
251–260 of 426 posts
Re: AI subscriptions are a ticking time bomb for enterprise
#252Re: AI subscriptions are a ticking time bomb for enterprise
#253Wasn't this the same thing when enterprises started using cloud computing? Did the bomb explode for them?
Re: AI subscriptions are a ticking time bomb for enterprise
#254Earlier quoted context omitted.
Do we know they are making a profit though? They could be subsidizing use to build market share the same way. They might not have billions, but at the volumes they are selling maybe they’ve got the cash to do it. Even if they are “profitable” how many Uber drivers are “profitable” because they aren’t correctly calculating asset depreciation. Maybe these guys are doing the same thing. Maybe it’s a lot of people who al…
> subsidizing use to build market share the same way To an extent maybe, but that market is almost entirely commoditized already. Besides Cerebras and maybe Groq (which already charge a slight premium) all the other providers are more less interchangeable. > Maybe it’s a lot of people who already had GPUs for crypto mining I’m not sure the type of GPUs that were most popular for crypto are at all useful for LLMs?
If there’s a few providers subsidizing, that’s the price ceiling. Everyone who wants to compete has to subsidize.
Now if this market had been operating for years, I’d say that it’s likely all these companies are profitable or close to it. But the market is so new and there’s so much hype, I find it very plausible that none of these guys are making a profit and they all hope to just hang in until all the subsidies go away.
> I’m not sure the type of GPUs that were most popular for crypto are at all useful for LLMs?
There’s some overlap. I’ve definitely read about people repurposing.
Re: AI subscriptions are a ticking time bomb for enterprise
#255Earlier quoted context omitted.
Isn't that always the case in the early stages of new technology adoption? It becomes less and less true as the new technology becomes more and more integrated. In the first few years after electric motors became a thing, one could have said the same thing. We would have just gone back to steam. If you tried to "do without them" now, society would collapse. So the question is not if we can do without them now, it's i…
The current LLM hype started, what, 5 years ago? It's an industry throwing billions of dollars (and teasing at the word trillions) around. It's had super bowl ads. It's a technology that's being mandated in corporate offices. It's basically the only thing the tech world ever talks about anymore. It's sucked all the air out of the room and occupies the whole stage. Just how "early stage" is that, and how much more int…
AI won't be "integrated" until something similar happens, and new businesses etc. are formed that take advantage of it in a way that can't simply be reversed to the old, pre-AI paradigm. I don't know what that will look like, but someone is going to figure it out and make successful companies with entirely new paradigms that are only made possible by AI.
At some point, every single factory was designed for electric motors, and going back became unthinkable.
-edit- also, the idea that a 5 year old tech that is still rapidly changing and developing deserves quotation marks around "new technology" is hilarious to me.
Re: AI subscriptions are a ticking time bomb for enterprise
#256Earlier quoted context omitted.
also, it's very much possible that the chinese companies get heavy investments from the state. Since it's very hard to get this info we have no idea wether they really make a profit or not.
The R&D is of course subsidized but a lot/most(?) of these inference providers are not Chinese
Re: AI subscriptions are a ticking time bomb for enterprise
#257[flagged]
"It costs OpenAI less money to serve GPT-5.5 than GPT-4." does it though? do you have the numbers? Or you just making stuff up?
Input: $30 / 1M tokens
Output: $60 / 1M tokens
GPT-5.5:
Input: $5 / 1M tokens
Output: $30 / 1M tokens
Costs have been reducing by over 5x year over year. Inference cost concern is mostly performative.
https://simianwords.bearblog.dev/conclusive-proofs-that-llm-...
Edit: can't reply but companies aren't selling inference at loss. In the blog post I point to third party hosting of open models like Deepseek which are also going down. They are not VC backed.
I also point to Gemma 31B which you can run on your laptop today that beats most models from 2024.
Re: AI subscriptions are a ticking time bomb for enterprise
#258Every AI subscription is a ticking time bomb for the frontier provider; within a few years we will be running local models as good as today’s frontier models with almost no cost burden. The floor will fall out of the enterprise market for all the frontier companies.
>within a few years Eventually, we'll see. Frontier models still need some pretty serious hardware which will slowly come down in cost. Smaller models are becoming more capable, which will presumably continue to improve. I think there's still a pretty big gap, though. Claude estimates Opus 4.6 and GLM-5 need about 1.5Ti VRAM. It puts gpt-5.5 around 3-6Ti of VRAM. That's 8x Nvidia H200 @ ~$30k USD each. Still need som…
Re: AI subscriptions are a ticking time bomb for enterprise
#259Even if they are momentarily losing money it’s important to note the value add they are providing. If you increase the price, the value is still astronomical in comparison. Companies need to find a way to leverage local models in tandem with frontier models to offset the costs. It’s all about targeting specific workloads with the appropriate AI. These tools are not sentient beings they are tools that need to be prope…
Re: AI subscriptions are a ticking time bomb for enterprise
#260Earlier quoted context omitted.
Capex, opex, quality, and volume are tricky things to balance. On balance, pc/mobile are cheaper to operate than equivalent cloud and on prem deployments. It’s not unreasonable to suppose that in 2 years time an opus 5 quality model will be etched into silicon for high performance local inference. Then you just upgrade your model every 2-3 years by upgrading your hardware.
I haven't been following anyone baking models into ASICs, is it not still necessary to pack just as many transistors onto a chip, whether it's an NPU or GPU, ASIC or not you still need to hold hundreds of gigabytes in memory, so how is it cheaper to bake it onto custom silicon than running it on commodity VRAM? (Asking because I don't know!)
Is an example startup in this area claiming 16k tok/s on an asic for llama 8b. Qwen has a 27b model at opus 4.5 quality.