Live data from Hacker News

GLM 5.2 and the coming AI margin collapse

martinalderson.com

251–260 of 495 posts

Re: GLM 5.2 and the coming AI margin collapse

#251
post #5

I'm not convinced raw costs matter: 1. Compute costs collapsed since the advent of Cloud and yet hyperscalers still have fat margins. 2. Many open source office suites exist yet none compete with the ubiquity of gsuite or office. GitHub, Slack are similar examples. 3. Both Windows and macOS dominate the home desktop space despite free alternatives existing for a long time. 4. Many formerly open source infrastructure…

For 2 and 3, office software and OSs have strong network effects and up-stack effects, just like CPU instruction sets. Also, I'm sorry, but OSS office suites compete with Office and GSuite the way grocery store frozen pizza competes with Domino's and Papa John's. Quality and completeness of execution matter a lot in that category.

Lol, your example of the other side of grocery store pizza is Domino’s and Papa John’s?

I guess in offices where M$ products are used the people there think mmm yumm dominos and hold up their noses at digiornos lol.

Re: GLM 5.2 and the coming AI margin collapse

#252
post #59
post #5

I'm not convinced raw costs matter: 1. Compute costs collapsed since the advent of Cloud and yet hyperscalers still have fat margins. 2. Many open source office suites exist yet none compete with the ubiquity of gsuite or office. GitHub, Slack are similar examples. 3. Both Windows and macOS dominate the home desktop space despite free alternatives existing for a long time. 4. Many formerly open source infrastructure…

> but I don't see any historical analogues. The losers are quickly forgotten. Palm, Blackberry, AOL, MySpace. Yahoo, etc. Software gets replaced all the time too, you even listed one and didn't realize. 15 years ago you'd call office irreplaceable, now you have to add gsuite to the mix, in 15 years there might be others. I know people that have never had office installed on their PC and use spreadsheets daily. > It s…

Because nobody’s paying for tokens, they’re paying for monthly plans and right now those are still a better bargain.

Re: GLM 5.2 and the coming AI margin collapse

#253
post #242
post #224

Earlier quoted context omitted.

> The next one is likely to be a doozy The US, EU, China are teetering on the edge of a crisis. Russia is well on its way. I feel like 2008 was just a warmup to what may be coming.

The EU is not "teetering on the edge of a crisis", why do you say that?

Not OP but, EU economy is being squeezed by China on the industrial and tech front, and by the US on the innovation/startup front. It is clear EU is no longer at the technological/economic frontier like it used to be 10-20 years ago. At the same time there are serious demographic, budgetary and political challenges all across the continent. Dragi's report covers some of these. It feels like the whole system might fall into a crisis soon if measures are not taken

Re: GLM 5.2 and the coming AI margin collapse

#254
I am confused by this article landing on front page, it does not seem to contain any new insights but besides mostly reading ok falls apart when mixing up model and harness comparison in an amateurish way. Why would the subpar search tool zai provides be relevant for comparing models? They did not even mention the >capability< to use said tools but talk about MCP/ search providers as if thats not an implementation detail.

Re: GLM 5.2 and the coming AI margin collapse

#255

Earlier quoted context omitted.

The current top comment in https://lobste.rs/s/ua1gxl/glm_5_2_coming_ai_margin_collapse correctly zoomed into cached input tokens, but landed on the opposite conclusion: > That is, for your $100/month fee, you get $3600 equivalent of API usage. This is presumably because Anthropic has figured out some clever things to do with model routing and input caching, and also can subsidize with investor money and take a hit o…

While we are all speculating, Boris kindly provided some guidance in https://news.ycombinator.com/item?id=47880089 > The challenge is: when you let a session idle for >1 hour, when you come back to it and send a prompt, it will be a full cache miss, all N messages. We noticed that this corner case led to outsized token costs for users. In an extreme case, if you had 900k tokens in your context window, then idled for…

I wouldn't be too fixated on the specific numbers in that post.

Anthropic was extremely capacity constrained at that point. They still are but not to that extent.

I'd note that OpenAI offers 24 hour caching. I'd be surprised if Anthropic hasn't optimised their caching for Claude code too.

SemiAnalysis recently posted that their actual Opus usage works out at $0.99 because of caching.

The principles remain though.

Re: GLM 5.2 and the coming AI margin collapse

#256

Earlier quoted context omitted.

The $5 is so they can see if open weights models are worth using, not so they can use it for a month. (Which you can't; The quota runs out way sooner than a month for any serious usage. Still worth the price of entry.)

If you use DeepSeek v4 Flash as a daily driver, with an occasional usage of DeepSeek V4 Pro and Glam 5.2 when necessary, the monthly quota practically never runs out.

Getting the pay-as-you-go plan from DeepSeek is also a good alternative. When motivation strikes you never get slowed down by quota, and it's cheap enough that even with mostly DeepSeek V4 Pro it's price-competitive with a $5/month subscription. Depending on how bursty your usage pattern is it might even be cheaper

Re: GLM 5.2 and the coming AI margin collapse

#257
post #224

Earlier quoted context omitted.

Somewhere else in the comments here, someone else remarked "Individuals perhaps [move to the new models], but not organizations." That's illustrative. The mechanism by which organizations are forced to update their technology, move to more competitive suppliers, and cut costs is a recession. In one, every business that doesn't do so goes bankrupt, and what's left are the more efficient businesses that have adopted te…

> The next one is likely to be a doozy The US, EU, China are teetering on the edge of a crisis. Russia is well on its way. I feel like 2008 was just a warmup to what may be coming.

> Russia is well on its way.

Including the warmongering angry midget next to the US, EU, and China is funny. Russia's economy, before they decided to shoot themselves in the face, was the size of the Netherlands. Whether they are in a recession or not is irrelevant to anyone but them.

Re: GLM 5.2 and the coming AI margin collapse

#258

Earlier quoted context omitted.

I agree, but there is prestige to consider. Many people are motivated to buy the best, even if it's much more expensive. "We're building a mission critical application here. Sure the API costs are much higher, but it's worth it."

I could spend $1000/day with Fable and it would be worth it. It has much deeper systems thinking, enabling me to trust it to follow my instructions and not fuck things up. 1. That confidence and quality is worth the price. 2. We're accelerating at lightning speed now. If you don't spend, someone else will and they'll eat your cake. We're nearing the point where you could spin up an entire YC startup in a day. That ch…

> We're accelerating at lightning speed now.

Accelerating how much slop you can output? A better model will still produce slop for your feature factory that pumps out software which nobody is interested in buying.

Re: GLM 5.2 and the coming AI margin collapse

#259
post #5

I'm not convinced raw costs matter: 1. Compute costs collapsed since the advent of Cloud and yet hyperscalers still have fat margins. 2. Many open source office suites exist yet none compete with the ubiquity of gsuite or office. GitHub, Slack are similar examples. 3. Both Windows and macOS dominate the home desktop space despite free alternatives existing for a long time. 4. Many formerly open source infrastructure…

Operating systems and office suites (like windows and word) have big network effects and high switching costs.

Much less with llm chatbots/coding tools.

Re: GLM 5.2 and the coming AI margin collapse

#260

Earlier quoted context omitted.

If you use DeepSeek v4 Flash as a daily driver, with an occasional usage of DeepSeek V4 Pro and Glam 5.2 when necessary, the monthly quota practically never runs out.

Getting the pay-as-you-go plan from DeepSeek is also a good alternative. When motivation strikes you never get slowed down by quota, and it's cheap enough that even with mostly DeepSeek V4 Pro it's price-competitive with a $5/month subscription. Depending on how bursty your usage pattern is it might even be cheaper

True, but OpenCode Go gives 6x tokens on Flash and 1.5x tokens on DeepSeek Pro. After exhausting the monthly quota, Flash price is the same as directly from DeepSeek, while Pro is 4x pricier.
Post reply on HN