GLM 5.2 and the coming AI margin collapse
11–20 of 495 posts
Re: GLM 5.2 and the coming AI margin collapse
#12Seems like a pretty pointless post that still centers around output tokens. In agentic coding, cached input tokens is 90% of the API "cost". It doesn't require GPU compute, and DeepSeek has shown that it can be done 50~100x cheaper with MLA/CSA/HCA, and a whole bunch of disks. This should collapse the margin.
Re: GLM 5.2 and the coming AI margin collapse
#13If your using pure API ... providers like neuralwatt cut that cost down even more by using energy as the actual cost. So GLM 5.2 is more expensive then GLM 5.1 on their service (those thinking tokens), compared to API costs, its dirt cheap. And way more tokens then the zai subscription delivers.
We are seeing a move towards more realistic pricing on actual consumption based usage. Be it DeepSeek, Xiaomi (MiMo), or zai's GLM via neuralwatt.
The main issue facing subscriptions a-la-carte usage, is that a lot of the heavy hitters really drain the resources. And that as a business model can not survive without ...
a) increasing the prices. b) everything goes to actual token/energy usage based billing but with more realistic pricing, and not the bloated API prices that are focused on companies.
We shall see what the future holds but things will change.
Re: GLM 5.2 and the coming AI margin collapse
#14Seems like a pretty pointless post that still centers around output tokens. In agentic coding, cached input tokens is 90% of the API "cost". It doesn't require GPU compute, and DeepSeek has shown that it can be done 50~100x cheaper with MLA/CSA/HCA, and a whole bunch of disks. This should collapse the margin.
> That is, for your $100/month fee, you get $3600 equivalent of API usage. This is presumably because Anthropic has figured out some clever things to do with model routing and input caching, and also can subsidize with investor money and take a hit on their operating margins.
My take: this is exactly what Anthropic wants everyone to think. In reality, 90% of that $3600 are for cached input tokens, that can be made to cost next to nothing, as shown by DeepSeek.
Re: GLM 5.2 and the coming AI margin collapse
#15I'm not convinced raw costs matter: 1. Compute costs collapsed since the advent of Cloud and yet hyperscalers still have fat margins. 2. Many open source office suites exist yet none compete with the ubiquity of gsuite or office. GitHub, Slack are similar examples. 3. Both Windows and macOS dominate the home desktop space despite free alternatives existing for a long time. 4. Many formerly open source infrastructure…
The target audience is different. Coding is mainly a trade of the tech savvy, who like many on r/localllama users do not hesitate to deply on 16GB Vram gpus. Even if so, it is estimated that within 2 years we will be able to run Claude 4.8 on consumer hardware give the rate of improvement of open-weight LLMs, which will put more financial pressure on "paid" labs. It's just a matter of rate of improvement which is shr…
Re: GLM 5.2 and the coming AI margin collapse
#16I also think we’ll approach a point where increasing intelligence is not really going to suddenly improve most work tasks. I bet that’s already happened actually. We’re oohing and aahing about models, when the ones a few versions ago did a good enough job for most of the dumb coding, etc we do
Requirements evolve in use, and fully hands-off LLMs simply cannot be trusted to only change the things you ask them to change, so I don't think it's likely that products will, in the main, be developed that way.
And if you don't need that fully-hands-off stuff, then the models that run on at least reasonably modest desktop hardware are surprisingly close to being enough.
Re: GLM 5.2 and the coming AI margin collapse
#17I'm not convinced raw costs matter: 1. Compute costs collapsed since the advent of Cloud and yet hyperscalers still have fat margins. 2. Many open source office suites exist yet none compete with the ubiquity of gsuite or office. GitHub, Slack are similar examples. 3. Both Windows and macOS dominate the home desktop space despite free alternatives existing for a long time. 4. Many formerly open source infrastructure…
And it's clear neither of the big two can deliver anything close to a service guarantee.
Re: GLM 5.2 and the coming AI margin collapse
#18This is why Google will win the race over most of its competitors. They own search.
Re: GLM 5.2 and the coming AI margin collapse
#19Re: GLM 5.2 and the coming AI margin collapse
#20> the least understood upcoming shift in AI economics. Then proceeds to talk about something in the AI news every day. Hey, did you guys hear? Open source models are cheaper and their quality is increasing! So, first, by no measure is GLM5.2 as good as Opus. Second, yes, open source models will put pressure on margins...eventually. Everyone knows that. But do you think today's AI business model is the same as tomorro…
Depends what you do. Complex tasks, poorly-defined tasks, sure. For relatively simple tasks, though, or very well-defined tasks, it's just as good and usually a lot faster. It also has a more neutral character and is somewhat less adversarial than Opus. (Opus is always "Let me push back on that..." whereas GLM is "sir, yes sir!") I use both and I appreciate both. If Opus disappeared tomorrow, though, I wouldn't cry -- I'd be able to adapt to a GLM-5.2-only life real quick.