Earlier quoted context omitted.
> MLA/CSA/HCA Aren’t these techniques all “lossy” compression, and one of the reasons people complain about loss in quality as the context size grows larger?
Indeed they are all lossy. Not sure how much they contribute to the quality loss in long context though. I got a 700k session with DSV4 Pro (official API), and the model was still coherent and didn't make any tool call error.
GLM 5.2 and the coming AI margin collapse
101–110 of 495 posts
Re: GLM 5.2 and the coming AI margin collapse
#102I'm not convinced raw costs matter: 1. Compute costs collapsed since the advent of Cloud and yet hyperscalers still have fat margins. 2. Many open source office suites exist yet none compete with the ubiquity of gsuite or office. GitHub, Slack are similar examples. 3. Both Windows and macOS dominate the home desktop space despite free alternatives existing for a long time. 4. Many formerly open source infrastructure…
Unlike all your examples, switching out an LLM is both cheap an easy. So easy that every 3 months or so new models are released and people grab them and start using them. The UX is the same regardless the provider. You send in a prompt, it spits back an answer. In all your other cases, the cost to switch is losing support and a difficult transition period. But in the case of LLMs, there was no support to begin with.…
For now
Re: GLM 5.2 and the coming AI margin collapse
#103I had GLM 5.2 do the same, and it performed exceptionally better, but when it got stuck on something it would be trial and error mode going forward and have zero foresight for future issues that might occur due to fixes it was trying. the model severally lacks prompt understanding, and testing .
Re: GLM 5.2 and the coming AI margin collapse
#104Earlier quoted context omitted.
They don't just need healthy margins, they need to make back almost a trillion dollars in a couple of years. Comparing that to elastic search and redis doesn't make much sense. Hyperscalers work because it actually has value compared to free offerings and because of the absolutely massive cost of switching providers. Similar with Windows and macOS. Extremely high cost of switching to something different, if possible…
The companies don't necessarily need to make back $1T, the investors do, and those investors don't require $1T in profit to do so, they need an asset worth $1T. Considering leaks suggest Anthropic's ARR would be $47B, that'd be a 20x valuation, but it wouldn't shock me if Anthropic doubles their revenue in the next year or two, in which a 10x revenue could easily support a $1T valuation, and boom there's your ROI, bu…
Re: GLM 5.2 and the coming AI margin collapse
#1051. There will be no moat around frontier AI models in the future. China is going to make sure that happens. It's a national security interest for them. DeepSeek was the first shot across the bow for that but it won't end with them. There are other labs and there are non-Chinese actors too. The stratospheric valuations depend on there being that moat; and
2. Nobody seems to be considering what the next generation of AI hardware is going to do with current hyperscalar investments. We're about to go through this with the B100/200 move to R100/200 but a lot of the investments are probably slated for that next-gen. But what about 3 years from now when the hypothetical X100/200 comes out and doubles FLOPS and halves performance-per-watt. What will that do to existing investments? Some people are delusional and think that they'll get 10 years out of GPUs when 10 year old GPUs (eg V100) are sold for scrap and 5 year old GPUs (A100) cannot run DeepSeek v4 Pro. And people think the A100 is going to get another 5 years of use? No; and
3. Local LLMs are coming for remote usage. You can buy a 5090 PC for less than $5000 currently but you're limited to 32GB of VRAM, which will comfortably run 31B models but nothing really larger. Go to $12-13k to upgrade to an RTX 8000 Pro and you have 96GB of VRAM, which will run larger models (but certainly not, say, DS v4 Pro or even Flash). You have shared video memory products rapidly coming from NVidia's aggressive market segmentation. Things like Strix Halo and DGX Spark have severe limits on memory bandwidth (But what will this local hardware look like in 2-3 years? I think people will be shocked at how much better it will be with the Apple M7 Pro/Max generation (2028 expected) and the RTX 6000 cards at that time although I fully expect NVidia consumer GPUs to still top out at 32GB of VRAM to maintain that segmentation. And I look forward to what the next generation of the AMD Ryzen AI Halo platform will look like if they really try.
All of this adds up to these three companies needing to cash out before the music stops (IMHO).
Re: GLM 5.2 and the coming AI margin collapse
#106The fact that these Chinese models are getting close to “Opus-grade” despite costing 6x-8x less is huge. As the token bills start to come in, those economics will be harder to ignore (regardless of the origin of the LLM); especially as there will be many CIOs sweating over their quick and costly AI initiatives showing little ROI. My hope is that the EU also steps up their own competition in the frontier model space s…
Re: GLM 5.2 and the coming AI margin collapse
#107Earlier quoted context omitted.
Unlike all your examples, switching out an LLM is both cheap an easy. So easy that every 3 months or so new models are released and people grab them and start using them. The UX is the same regardless the provider. You send in a prompt, it spits back an answer. In all your other cases, the cost to switch is losing support and a difficult transition period. But in the case of LLMs, there was no support to begin with.…
If you're developing on top of LLM APIs directly, this is definitely not true. There are differences in how context caching works, in what's available through native harnesses, the types of tools you're fine-tuned on (GPT uses apply_patch while Claude uses edit, with different formats), the API surface (Agents SDK, Responses API, Managed Agents), cost structures, and best-practice guidance all around. Not to mention…
If anything, you're better off supporting multiple LLMs as backup because most model providers have been so inconsistent with working all the time
Re: GLM 5.2 and the coming AI margin collapse
#108Earlier quoted context omitted.
Unlike all your examples, switching out an LLM is both cheap an easy. So easy that every 3 months or so new models are released and people grab them and start using them. The UX is the same regardless the provider. You send in a prompt, it spits back an answer. In all your other cases, the cost to switch is losing support and a difficult transition period. But in the case of LLMs, there was no support to begin with.…
I don’t exactly see orgs lining up to switch (and train) their employees between claude desktop and codex and whatever copilot is doing. There’s probably some inertia to those harnesses/integrations on top of the llms themselves.
There's barely any moat. All the data is with connectors, memory is near useless
Re: GLM 5.2 and the coming AI margin collapse
#109Earlier quoted context omitted.
Unlike all your examples, switching out an LLM is both cheap an easy. So easy that every 3 months or so new models are released and people grab them and start using them. The UX is the same regardless the provider. You send in a prompt, it spits back an answer. In all your other cases, the cost to switch is losing support and a difficult transition period. But in the case of LLMs, there was no support to begin with.…
I don’t exactly see orgs lining up to switch (and train) their employees between claude desktop and codex and whatever copilot is doing. There’s probably some inertia to those harnesses/integrations on top of the llms themselves.
Re: GLM 5.2 and the coming AI margin collapse
#110I think the profits depend on how well they manage their fleet purchases (or possible sub-leasing?) to get high utilization without overloading or idle racks. Because accelerators like H200, B300 etc. are highly parallel and designed to run like 200 or maybe 300 sequences at once (depends on the model, just guessing). I assume they finance the hardware and that cost per device or rack is the same whether each unit is…