Live data from Hacker News

GLM 5.2 and the coming AI margin collapse

martinalderson.com

301–310 of 495 posts

Re: GLM 5.2 and the coming AI margin collapse

#301
The problem is that the more AI eats labor, the more you can hike the costs until you pretty much can match the salary of the workers you replaced, with some margin, enough for the user to accept the cost. That’s what will happen in the next decade IMHO. price = base expense + what user accepts to pay

Re: GLM 5.2 and the coming AI margin collapse

#302
post #5

I'm not convinced raw costs matter: 1. Compute costs collapsed since the advent of Cloud and yet hyperscalers still have fat margins. 2. Many open source office suites exist yet none compete with the ubiquity of gsuite or office. GitHub, Slack are similar examples. 3. Both Windows and macOS dominate the home desktop space despite free alternatives existing for a long time. 4. Many formerly open source infrastructure…

[dead]

Re: GLM 5.2 and the coming AI margin collapse

#303

I'll agree but from the other direction. AI continues to absorb my job as a senior systems software engineer (c/c++) and after a couple months I've only spent a few hundred dollars using gpt-5.5/5.6 and codex. I have no idea what people are doing to burn so many tokens but for me this is laughably cheap and every day I discover new capabilities. I don't care if costs go up or down, it's so cheap for what I get that I…

These are probably mostly the enterprise customers - they may use the same amount of tokens as you do, but they have to pay the API price. From my experience the API is significantly more costly. We had one user ask for and receive usage credits on Claude, the bill the next day was to the tune of $400.

That's around 100k/year if used at the same rate for every workday. So the question becomes: does it make your engineer X% more productive, where X is some multiplier based on their salary? There are some software engineers out for sure who are expensive enough that this is worth it.

Re: GLM 5.2 and the coming AI margin collapse

#304
post #5

I'm not convinced raw costs matter: 1. Compute costs collapsed since the advent of Cloud and yet hyperscalers still have fat margins. 2. Many open source office suites exist yet none compete with the ubiquity of gsuite or office. GitHub, Slack are similar examples. 3. Both Windows and macOS dominate the home desktop space despite free alternatives existing for a long time. 4. Many formerly open source infrastructure…

good observation! however, I'd argue it's the distribution channel and installation friction did the job in the cases you mentioned.

given how easy it's to replace LLM API in claude code, and how easy it is to write a claude code clone with itself (Fable is pretty good!), the collapse is coming.

Re: GLM 5.2 and the coming AI margin collapse

#305
post #292

Regarding the lack of vision part, if you are using Claude or opencode, I've made a skill[1] that let's you talk with any models in Claude/opencode mid-session. You ask "Have claude opus to look at this PDF for a second opinion" during a session of claude with GLM5.2 or opencode with GLM5.2 It doesn't need to pass whole conversation history as context (unlike /model), you can ask follow up to that forked model (which…

FWIW I've seen subagents remain open for followups on latest Claude.

Re: GLM 5.2 and the coming AI margin collapse

#306

I'll agree but from the other direction. AI continues to absorb my job as a senior systems software engineer (c/c++) and after a couple months I've only spent a few hundred dollars using gpt-5.5/5.6 and codex. I have no idea what people are doing to burn so many tokens but for me this is laughably cheap and every day I discover new capabilities. I don't care if costs go up or down, it's so cheap for what I get that I…

It seems like you're basing your spend on the subsidized consumer subscriptions. The equivalent API costs for these subscriptions is usually 12-20x. Any overages (hourly/weekly/model) on these plans gets billed at rack API costs. Its not practical to expect these subsidies to last for very long.

Agentic workloads are the most batch friendly, latency insensitive, geography insensitive, migration insensitive tokens that a big lab ever sells. In the ads business such inventory is called "remnant". The sausage is made of whatever is left over when the choice cuts have been removed.

This talking point from Anthropic that Claude Code sitting in a Ralph Loop is burning top sirloin interactive session tokens is bad faith hogwash and it only flies because most everyone who has run this shit at scale either already works there, sells them hardware, or hopes to be an acquisition target.

I'm none of those things, so I'm happy to tell you they're lying. I know, it's hard to swallow, but it turns out Altman and Amodei are occasionally full of shit.

Re: GLM 5.2 and the coming AI margin collapse

#307

Earlier quoted context omitted.

Because we pay retail consumer prices (subscriptions). Those same tokens cost many thousands on enterprise billing:/

What makes you say that the parent you’re replying to uses a consumer subscription? Sounds to me like he’s using it for work.

Because if this person was paying enterprise for tokens, they:

1. wouldn't write stuff like "I've only spent a few hundred dollars using gpt-5.5/5.6 and codex"

2. wouldn't think tokens are cheap

Re: GLM 5.2 and the coming AI margin collapse

#308
post #271

Earlier quoted context omitted.

Factor in that Europe is powerful partly because of its unity and there are many forces trying to undermine that.

Europe became powerful before it was unified, and ever since the creation of the EU it's been becoming less and less important on the world stage.

The motivation that the USA entered WWII was not because they were generous, but because the 3rd Reich was effectively becoming a big European nation, so they had to do something to avoid it. A unified Europe is a thread to the USA and Russia and maybe somebody else too.

Re: GLM 5.2 and the coming AI margin collapse

#310

Earlier quoted context omitted.

AWS already supports Llama and GLM in its Bedrock service for hosted models. They’re much cheaper to run, eg, Llama 3.3 Instruct 70B is 5-10x cheaper than Sonnet 5. https://aws.amazon.com/bedrock/pricing/ Say you have 20% of usecases that require the more expensive model — but in 80% you could just use Llama instead of Sonnet (eg, for basic queries of a document). That saves 80% of that 80%, or 65% of your total bill…

Bedrock is really out of date with the models it offers, to the extent that I'm not sure they even have plans to update what's on there now they have the deal with Anthropic. They're still offering Qwen 3 , not even 3.5 and certainly not 3.6. GLM 5 is the newest z.AI model they have, when it's 5.2 that would be the one to worry Sonnet. There are some ok models on there (Qwen 3 Coder Next is usable and fast, for insta…

Because these models are not going to stand up legally
Post reply on HN