Live data from Hacker News

GLM 5.2 and the coming AI margin collapse

martinalderson.com

231–240 of 495 posts

Re: GLM 5.2 and the coming AI margin collapse

#231

> the least understood upcoming shift in AI economics. Then proceeds to talk about something in the AI news every day. Hey, did you guys hear? Open source models are cheaper and their quality is increasing! So, first, by no measure is GLM5.2 as good as Opus. Second, yes, open source models will put pressure on margins...eventually. Everyone knows that. But do you think today's AI business model is the same as tomorro…

> So, first, by no measure is GLM5.2 as good as Opus.

That's an opinion many will disagree with. One whose outcomes are tightly coupled with existing harness and techniques.

In my real life usage Opus 4.7 and 4.8 have been increasingly unhelpful compared to 4.6 in behaving as assistants.

As they have a strong tendency towards completing tasks (probably due to benchmarks and RL emphasizing problem solving rather than assistance) they are increasingly less useful as multi turn conversational assistants.

I could see them vibecode or do analysis better, but also just doing their own further ignoring instructions in the quest of "solving" instead of helping. Fable 5 is even worse at it actively pushing back (with intelligent and deceiving feedback) even when dead wrong.

GLM seems to suffer less of this.

Re: GLM 5.2 and the coming AI margin collapse

#232
post #5

I'm not convinced raw costs matter: 1. Compute costs collapsed since the advent of Cloud and yet hyperscalers still have fat margins. 2. Many open source office suites exist yet none compete with the ubiquity of gsuite or office. GitHub, Slack are similar examples. 3. Both Windows and macOS dominate the home desktop space despite free alternatives existing for a long time. 4. Many formerly open source infrastructure…

Examples two and three largely persist due to massive vendor lock in after the vendor has done enough work to capture market share, but that does not seem to be the case for AI labs to my knowledge

Re: GLM 5.2 and the coming AI margin collapse

#233
post #105

I think OpenAI, Anthropic and SpaceX are going to envy the dinosaurs because there's not asteroid coming for them, there's three: 1. There will be no moat around frontier AI models in the future. China is going to make sure that happens. It's a national security interest for them. DeepSeek was the first shot across the bow for that but it won't end with them. There are other labs and there are non-Chinese actors too.…

On 1 and 3, the obvious move is to shift the bulk of the harness behind a new API that's not based on raw LLM access. Then they get to hide secret sauce behind that API and all three go from commodity to premium while simultaneously being able to try out whatever tricks they can get away with to reduce their own inference costs. I'm almost surprised this hasn't happened already.

Re: GLM 5.2 and the coming AI margin collapse

#234

Earlier quoted context omitted.

Unlike all your examples, switching out an LLM is both cheap an easy. So easy that every 3 months or so new models are released and people grab them and start using them. The UX is the same regardless the provider. You send in a prompt, it spits back an answer. In all your other cases, the cost to switch is losing support and a difficult transition period. But in the case of LLMs, there was no support to begin with.…

Switching an agent harness is more difficult, especially on the enterprise/teams level. Once your team gets settled with Claude teams, cowork, and the various plugins, it’s going to be a pain in the butt to switch.

I made my own. It's true I don't want to switch to another harness now.

But switching models is just a command.

Re: GLM 5.2 and the coming AI margin collapse

#235

Earlier quoted context omitted.

I'd like to understand this please. Why would the 1M context be kept in VRAM if you're using DSV4 Pro through the API? Or did you refer to different sessions?

Different sessions. With https://github.com/fairydreaming/llama.cpp/tree/dsv4 , 1M context with DSV4 Flash takes less than 6GB of VRAM. I can't run DSV4 Pro, but it should take less than 9GB of VRAM for 1M context, based on the numbers shared in https://arxiv.org/html/2606.19348v1 .

Thank you for the links/docs. I'm quite excited to try it myself.

Re: GLM 5.2 and the coming AI margin collapse

#236
Unlike the belief that frontier AI is expensive due to a high margin, and going to be expensive if there is no competition. My understanding is that, under certain circumstances (which is most likely true), the price will be driven down just because of profit seeking.

The frontier LLM labs run on a huge fixed cost and very low marginal cost. They need the economies of scale to make sense of the business (an incentive to expand their user base as large as possible). Imagine that you want to buy a few B300s to run GLM 5.2 and rent the service out to other people. How could this business be viable and sustainable in the first place? You need as many customers as possible. If you charge everyone $1000, you find fewer customers who can afford it. It rots the ROA if the servers are not utilized 100% (you would better buy less compute instead).

Also, the marginal cost for onboarding a new customer is low. And it's getting even lower when you have more customers. You wouldn't leave money on the table (especially for your competitors) if you want to maximize your profit.

By this logic, all frontier AI labs are incentivized to lower the price to maximize their customer base, profit, and ROA.

Re: GLM 5.2 and the coming AI margin collapse

#237
post #196

Earlier quoted context omitted.

> Just spend $5 on OpenCode Go $5 the first month , then price is doubled.

The $5 is so they can see if open weights models are worth using, not so they can use it for a month. (Which you can't; The quota runs out way sooner than a month for any serious usage. Still worth the price of entry.)

If you use DeepSeek v4 Flash as a daily driver, with an occasional usage of DeepSeek V4 Pro and Glam 5.2 when necessary, the monthly quota practically never runs out.

Re: GLM 5.2 and the coming AI margin collapse

#238
post #55

The fact that these Chinese models are getting close to “Opus-grade” despite costing 6x-8x less is huge. As the token bills start to come in, those economics will be harder to ignore (regardless of the origin of the LLM); especially as there will be many CIOs sweating over their quick and costly AI initiatives showing little ROI. My hope is that the EU also steps up their own competition in the frontier model space s…

they're not near opus at all, anyone using the models in a real working environment will tell you the same thing. on paper they have impressive benchmarks, but that's not realistic to actual use.

I think it depends on your use case. For my personal projects (a mix of webdev & some Rust desktop apps) it's honestly very close to Opus 4.8 (which I use in my day job).

I don't feel like I'm missing out after cancelling my personal Claude subscription, whereas I used to feel that way a few months ago.

Re: GLM 5.2 and the coming AI margin collapse

#239

I'll agree but from the other direction. AI continues to absorb my job as a senior systems software engineer (c/c++) and after a couple months I've only spent a few hundred dollars using gpt-5.5/5.6 and codex. I have no idea what people are doing to burn so many tokens but for me this is laughably cheap and every day I discover new capabilities. I don't care if costs go up or down, it's so cheap for what I get that I…

Because we pay retail consumer prices (subscriptions). Those same tokens cost many thousands on enterprise billing:/

What makes you say that the parent you’re replying to uses a consumer subscription? Sounds to me like he’s using it for work.

Re: GLM 5.2 and the coming AI margin collapse

#240

> I expect for most professional use the very lax terms around training and data retention will make this a difficult sell I couldn't care less whether a chinese or american company reads my crap code. I'm not working on state secrets but warehousing software for specific clients on a machine that has access to nothing but crap enterprise code.

This take may be frowned upon in HN, but I suspect it's much closer to the median sentiment.
Post reply on HN