This logic seems mad. If people only need SLMs then hyperscalers can also centrally host higher-efficiency models, and still gain efficiencies of scale and convenience over hosting locally.
The logic seems mad to me because SLMs can simply not hold as much information as an LLM. Maybe if you combine an SLM with a database (as a tool) then it could work, but someone should first prove that.
If this is true, the hyperscalers are toast
51–60 of 108 posts
Re: If this is true, the hyperscalers are toast
#52"If", sure. How many developers here don't see a difference between the latest LLMs and SLMs they can run on their own computer? I tried running a smaller model locally, and it's not usable for me. I know people like to "predict" things, so that if they happen they can then say "I am a visionary, I predicted it" and start their blog posts with "as I predicted long ago (because I am a visionary), ...". > The research…
>> I stopped counting the number of times "estimates" said that a market would absolutely explode, and it absolutely didn't. Those are in the business of being a broken clock. The lack of logic and risk management on this statement, is so strong, I hope humans are all quickly substituted by LLMs. Lets just do it and be done with it...
Why don't you go talk to your LLM instead of commenting here, then?
Re: If this is true, the hyperscalers are toast
#53Re: If this is true, the hyperscalers are toast
#54Earlier quoted context omitted.
> I tried running a smaller model locally, and it's not usable for me. If you have the hardware, a MacBook Pro for Qwen 3.6 35B A3B and Gemma 4 26B A4B for example, they are absolutely usable, both in terms of speed and quality. Anecdotally, I can use Qwen for day-to-day coding tasks in TS and Go, without hickups.
I'm unable to find a local model that comes close to the effectiveness of GPT models in Codex, and I have 96GB of VRAM available and tried every local model under the sun so far. Neither of those you mention I'd say are good enough for day to day software engineering for me, but I'm also really strict about code quality and iterate on what outputs agents give me a lot before I'm happy. With local models, this iterati…
This is becoming increasingly important to me. Super smart max reasoning frontier is fine if I leave it running overnight on some prepared set of clearly defined tasks, but when I want to work with the LLM, throughput really matters, and I'll go with a dumber model to get there.
At some point though, it's fast enough and any speed gains beyond that just makes me the bottleneck.
I also am seeing the smaller models gaining big strides lately, closing the gap on frontier models (still a decent sized gap though). I don't even run the small models like Qwen 3.8 27B locally. I just try them out in the cloud to see how they are progressing, and I'm definitely able to be productive.
Re: If this is true, the hyperscalers are toast
#55From what’s presented this seems to be the lower end of Q&A and reasoning tasks and not long horizon agentic work. I agree that the search engine replacement AI usage is something that can run anywhere (though it’s still better run in the cloud for speed, context length, sandboxing and convenience) but this isn’t the engine of AI growth. Also, the average consumer is not going to be running a local model until they a…
Apple or Nvidia, presumably.
Re: If this is true, the hyperscalers are toast
#56Then of course there is the economics of it. Do people prefer to spend $5000 upfront to get things done 5x slower, or would they rather pay $20 a month for that?
Re: If this is true, the hyperscalers are toast
#57Earlier quoted context omitted.
But do you really need a model that has the complete Duran Duran discography memorized and preloaded in RAM at all time?
That's a different question. Probably not. But: 1. Training a large model with lots of information, then stripping the "useless" information from that model to obtain a small model => nobody has shown this. 2. Training a small model, letting it use a database tool so it scores the same as a large model without database => nobody has shown this.
To go very small (thousands of rules) so that the reasoner can be understood by humans and proven sound - might be computationally quite difficult.
Re: If this is true, the hyperscalers are toast
#58From what’s presented this seems to be the lower end of Q&A and reasoning tasks and not long horizon agentic work. I agree that the search engine replacement AI usage is something that can run anywhere (though it’s still better run in the cloud for speed, context length, sandboxing and convenience) but this isn’t the engine of AI growth. Also, the average consumer is not going to be running a local model until they a…
> Also, the average consumer is not going to be running a local model until they are built into the hardware they already buy and when they are, who is supplying the weights? Apple or Nvidia, presumably.
Re: If this is true, the hyperscalers are toast
#59I don't think they are toast, I mean they will be in some trouble because all of them have fallen victim to fomo and started building out with so much debt for capacity that may or may not be needed nor achieve the returns that they want. I think there's a future where "personal software" meaning highly custom apps generated by an agent is a thing that doesn't mean everything will become that, same for local LLMs but…
1. The hyperscalers are in a positive reinforcement loop. Despite any suggestions to the contrary they keep getting bigger. And can, er, “influence” government policy/officials and anything else needed to keep it that way.
2. The frontier labs and their investors. Another self-fulfilling reinforcement loop. Witness the circular gymnastics among OAI/Anthropic, Microsoft/Amazon and Nvidia
3. Data. No-one believes that Zuckerberg and co are going to say “great, we can just run the models on devices we don’t own and stop the surveillance economy because, y’know, privacy matters and we really care about mental health”.
And then there’s data centre locations and “yeah but jobs” even though your power bills are going up, and “why run your own data centre Mrs CTO, let us do it for you and save all that capex and those pesky employees you need to do it”.
Don’t get me wrong: I’m rooting for local, open weight/source models. But “hey look they benchmark well” is an unhelpfully narrow basis to forecast the demise of central hyperscaler hosting.
Re: If this is true, the hyperscalers are toast
#60This logic seems mad. If people only need SLMs then hyperscalers can also centrally host higher-efficiency models, and still gain efficiencies of scale and convenience over hosting locally.
The valuations of the hyperscalars won't sustain just being more efficient than something you can run locally. There's a market there, but it's for margin on a commodity. They're priced for oligopoly on unique, premium products.
The compute capacity, however, is here to stay. There will be a seemingly endless line of people queued up to buy it at pennies on the dollar if this whole thing blows up financially. And they will use it!