It’s important that none of these entities can collude to price fix. Having China be the competitor ensures that. Basic microeconomics is still the easiest way to understand token economies. How is it not a competitive market (where profits go to zero?). Anything A or O does to keep more margin, any competitor can copy or choose to undercut, and undercutting has the benefit of collecting training data. So what is goi…
GLM 5.2 and the coming AI margin collapse
51–60 of 495 posts
Re: GLM 5.2 and the coming AI margin collapse
#52How long will that $4.40 rate persist? Until we know more about the real unit economics it will be damn near impossible to rely on steady inference costs or make them predictable at the enterprise level. Gonna be a wild ride for awhile.
Multiple providers (who need to make a profit) offer the same 4.40 rate for glm-5.2. It's not subsidized. Deepseek's 0.86 or whatever is likely subsidized but alternate providers offer it for a price comparable to glm-5.2.
Re: GLM 5.2 and the coming AI margin collapse
#53I hope cheaper inference eventually means faster speeds at the lower tiers. I don't want to settle for 100 t/s, but I don't want to pay $10 per prompt either
Which raises the question, which are the fastest frontier models? are the enterprise hosted Anthropic models faster than what Anthropic serves? Somehow no one talks about LLM speed.
When I've raised speeds about local inference I've been told 60-75 t/s is perfectly usable. It makes sense that people aren't talking about speed yet since you either already have a response fast enough to wait for, or you go do something else and check back in a few minutes.
I would love to wait for the latter type of tasks though, because those are typically the ones that require the most work from me to verify and I don't want my attention divided with multitasking.
Re: GLM 5.2 and the coming AI margin collapse
#54It’s important that none of these entities can collude to price fix. Having China be the competitor ensures that. Basic microeconomics is still the easiest way to understand token economies. How is it not a competitive market (where profits go to zero?). Anything A or O does to keep more margin, any competitor can copy or choose to undercut, and undercutting has the benefit of collecting training data. So what is goi…
Re: GLM 5.2 and the coming AI margin collapse
#55As the token bills start to come in, those economics will be harder to ignore (regardless of the origin of the LLM); especially as there will be many CIOs sweating over their quick and costly AI initiatives showing little ROI.
My hope is that the EU also steps up their own competition in the frontier model space so that it’s not just China v USA.
Re: GLM 5.2 and the coming AI margin collapse
#56I'm not convinced raw costs matter: 1. Compute costs collapsed since the advent of Cloud and yet hyperscalers still have fat margins. 2. Many open source office suites exist yet none compete with the ubiquity of gsuite or office. GitHub, Slack are similar examples. 3. Both Windows and macOS dominate the home desktop space despite free alternatives existing for a long time. 4. Many formerly open source infrastructure…
Those solutions have moats: 1. the cloud moat is mostly around talent really. Try finding people who can self host the alternatives to S3 et al at the HA and the scale the businesses need. Those alternatives are usually not free either, and each product might have its creator acquired (and the product cancelled) or similar. if you're a larger business then the data lock in becomes a moat: getting your data out of the…
Re: GLM 5.2 and the coming AI margin collapse
#57> It turns out that nearly every agentic session does a lot of web searching for looking up items This is why Google will win the race over most of its competitors. They own search.
I know Brave do this already. Not sure about DDG (I wonder if their agreement with Bing would allow it?)
Re: GLM 5.2 and the coming AI margin collapse
#58Earlier quoted context omitted.
Which raises the question, which are the fastest frontier models? are the enterprise hosted Anthropic models faster than what Anthropic serves? Somehow no one talks about LLM speed.
OAI has announced an upcoming 750tok/s 5.6 served through their cerebras acquisition
Re: GLM 5.2 and the coming AI margin collapse
#59I'm not convinced raw costs matter: 1. Compute costs collapsed since the advent of Cloud and yet hyperscalers still have fat margins. 2. Many open source office suites exist yet none compete with the ubiquity of gsuite or office. GitHub, Slack are similar examples. 3. Both Windows and macOS dominate the home desktop space despite free alternatives existing for a long time. 4. Many formerly open source infrastructure…
The losers are quickly forgotten. Palm, Blackberry, AOL, MySpace. Yahoo, etc.
Software gets replaced all the time too, you even listed one and didn't realize. 15 years ago you'd call office irreplaceable, now you have to add gsuite to the mix, in 15 years there might be others. I know people that have never had office installed on their PC and use spreadsheets daily.
> It seems that enterprises will pay top dollar for service guarantees, integration, and someone they can sue.
Of course. But why pay $25 per million tokens for sonnet when you can pay $3 for GLM? Both probably running on AWS/Azure/Etc. under some third party.