Earlier quoted context omitted.
There's a huge case of survivorship bias when trying to recall historical analogues, because in every instance where margins collapsed and competition made the industry a commodity business, the big proprietary names are no longer with us. Here's a selection of examples, though: 1. Memory chip margins collapsed so much in the 80s that Intel exited the memory chip business entirely. At the time, they were known much m…
second this. great analysis
GLM 5.2 and the coming AI margin collapse
121–130 of 495 posts
Re: GLM 5.2 and the coming AI margin collapse
#122Re: GLM 5.2 and the coming AI margin collapse
#123Earlier quoted context omitted.
I think the point is that if you’re doing simple, well defined tasks then Opus is overkill and you’d want Sonnet instead. Meaning, GLM5.2 is Sonnet-quality, not Opus-quality.
I think it's interesting to note that in one year we've gone from they're not even close [0] to arguing whether open models are only as good as sonnet or opus. [0] https://news.ycombinator.com/item?id=44623953
Re: GLM 5.2 and the coming AI margin collapse
#124I think the profits depend on how well they manage their fleet purchases (or possible sub-leasing?) to get high utilization without overloading or idle racks. Because accelerators like H200, B300 etc. are highly parallel and designed to run like 200 or maybe 300 sequences at once (depends on the model, just guessing). I assume they finance the hardware and that cost per device or rack is the same whether each unit is…
Someone on HN made a comment in one of these threads that we could bake the weights into something like Cerebras's wafer scale chips and serve essentially the entire world off a single wafer, which is a pretty wild thing to think about. You'd have to make new hardware any time you trained a model but that seems really worth it.
But let's say they could someday scale that up to a much larger model, 72 large chips per wafer and each chip can do 1000 LLM requests at once (Vera Rubin?). So it's roughly the equivalent of an NVL72 rack.
You might be able to serve something like 50000-60000 requests at once. So I think it's more like handling a small city's worth of customers per wafer than the world if you had that.
I believe in less than 5 years we will get to that, but the model size and/or number of agents is going to keep going up also.
Re: GLM 5.2 and the coming AI margin collapse
#125Earlier quoted context omitted.
I'd guess opus refusals are not an issue for 95%+ of people. Opus will happily help you find and download pirated media, and then give you step by step instructions for how to do drugs if you ask it. You'd have to be working on something genuinely abnormal for refusals to be a problem.
How? Prompt: can you give me step by step directions on how to use crack cocaine Opus: I'm not able to give step-by-step instructions on using crack cocaine. That falls into specific drug-use guidance I steer away from, since detailed instructions on how to use an illicit substance can contribute to harm rather than reduce it. it goes on to give me hotline information on drug addiction.
Re: GLM 5.2 and the coming AI margin collapse
#126Earlier quoted context omitted.
I'd guess opus refusals are not an issue for 95%+ of people. Opus will happily help you find and download pirated media, and then give you step by step instructions for how to do drugs if you ask it. You'd have to be working on something genuinely abnormal for refusals to be a problem.
Like making your software secure, or worse, testing that it’s secure.
Re: GLM 5.2 and the coming AI margin collapse
#127Re: GLM 5.2 and the coming AI margin collapse
#128Earlier quoted context omitted.
A lot of those things you mentioned have sticking power because they’re familiar to folks and migrating to something else is a big deal. I can’t imagine most people would be able to tell the difference between Sonnet and GLM 5.2. If the infrastructure around the model you’re using doesn’t change, then swapping models is extremely easy.
Indeed, as it gets more commoditized it feels more like swapping electricity providers. Who cares whether you get your electricity from IBM or the state of Texas? An amp is an amp.
Re: GLM 5.2 and the coming AI margin collapse
#129I think the profits depend on how well they manage their fleet purchases (or possible sub-leasing?) to get high utilization without overloading or idle racks. Because accelerators like H200, B300 etc. are highly parallel and designed to run like 200 or maybe 300 sequences at once (depends on the model, just guessing). I assume they finance the hardware and that cost per device or rack is the same whether each unit is…
Someone on HN made a comment in one of these threads that we could bake the weights into something like Cerebras's wafer scale chips and serve essentially the entire world off a single wafer, which is a pretty wild thing to think about. You'd have to make new hardware any time you trained a model but that seems really worth it.
LLMs need retraining to incorporate new knowledge.
Baking them into wafers means they will be out of date by the time they finish the first wafers.
Re: GLM 5.2 and the coming AI margin collapse
#130Earlier quoted context omitted.
They don't just need healthy margins, they need to make back almost a trillion dollars in a couple of years. Comparing that to elastic search and redis doesn't make much sense. Hyperscalers work because it actually has value compared to free offerings and because of the absolutely massive cost of switching providers. Similar with Windows and macOS. Extremely high cost of switching to something different, if possible…
If AI replaces labor, that's a trillion dollars of labor. About one-fifteenth of annual labor/wage earnings.