This logic seems mad. If people only need SLMs then hyperscalers can also centrally host higher-efficiency models, and still gain efficiencies of scale and convenience over hosting locally.
If this is true, the hyperscalers are toast
71–80 of 108 posts
Re: If this is true, the hyperscalers are toast
#72Earlier quoted context omitted.
The valuations of the hyperscalars won't sustain just being more efficient than something you can run locally. There's a market there, but it's for margin on a commodity. They're priced for oligopoly on unique, premium products.
Yes, it's all about the amount of money invested right now. Justifying that was always going to be tricky, and it still is. The compute capacity, however, is here to stay. There will be a seemingly endless line of people queued up to buy it at pennies on the dollar if this whole thing blows up financially. And they will use it!
Re: If this is true, the hyperscalers are toast
#73From what’s presented this seems to be the lower end of Q&A and reasoning tasks and not long horizon agentic work. I agree that the search engine replacement AI usage is something that can run anywhere (though it’s still better run in the cloud for speed, context length, sandboxing and convenience) but this isn’t the engine of AI growth. Also, the average consumer is not going to be running a local model until they a…
Says who? In the world of video decoding, H.265 is losing out to AV1 largely because it's not superior enough to H.264 to justify the expensive and complicated licensing.
Do you really think chipmakers are going to pay a 25c per unit tax to get a model that is 15% better?
Re: If this is true, the hyperscalers are toast
#74Earlier quoted context omitted.
Yes, it's all about the amount of money invested right now. Justifying that was always going to be tricky, and it still is. The compute capacity, however, is here to stay. There will be a seemingly endless line of people queued up to buy it at pennies on the dollar if this whole thing blows up financially. And they will use it!
> And they will use it! Just out of curiosity: Use it for what?
The demand is there and the market will find pricing that works.
Re: If this is true, the hyperscalers are toast
#75Earlier quoted context omitted.
> Also, the average consumer is not going to be running a local model until they are built into the hardware they already buy and when they are, who is supplying the weights? Apple or Nvidia, presumably.
Hardware yes, weights? lol
Re: If this is true, the hyperscalers are toast
#76Earlier quoted context omitted.
The valuations of the hyperscalars won't sustain just being more efficient than something you can run locally. There's a market there, but it's for margin on a commodity. They're priced for oligopoly on unique, premium products.
I very much don’t want to run it locally. I want the same one running somewhere else that I can interact with from all my devices. Look at something like Grok Bot. Nobody is going to run this locally. You can already self host almost anything, yet most people and businesses don’t.
I run models on my desktop and access them from a phone app on the go. Wireguard tunnel.
Responds fast and lets me kick off tasks or workflows via text or voice.
Re: If this is true, the hyperscalers are toast
#77Earlier quoted context omitted.
I very much don’t want to run it locally. I want the same one running somewhere else that I can interact with from all my devices. Look at something like Grok Bot. Nobody is going to run this locally. You can already self host almost anything, yet most people and businesses don’t.
Yes, but that's not a 10T business.
Re: If this is true, the hyperscalers are toast
#78Earlier quoted context omitted.
> I tried running a smaller model locally, and it's not usable for me. If you have the hardware, a MacBook Pro for Qwen 3.6 35B A3B and Gemma 4 26B A4B for example, they are absolutely usable, both in terms of speed and quality. Anecdotally, I can use Qwen for day-to-day coding tasks in TS and Go, without hickups.
I'm unable to find a local model that comes close to the effectiveness of GPT models in Codex, and I have 96GB of VRAM available and tried every local model under the sun so far. Neither of those you mention I'd say are good enough for day to day software engineering for me, but I'm also really strict about code quality and iterate on what outputs agents give me a lot before I'm happy. With local models, this iterati…
I agree with you there, the local models are not as capable as frontier remote LLMs. If you‘re used to letting fable run for an hour to do a novel or complex task, you‘re (probably?) not going to be happy with local LLMs. But, to me, for my daily work, they‘re still usable, and often surprisingly capable. YMMV
Re: If this is true, the hyperscalers are toast
#79One of the big things to think about is whether local LLMs will be things companies want to deploy. If you think of for e.g. some proprietary piece of software that wants to embed an LLM they've fine tuned or trained, they will want to make back some of their research cost right. So they are not going to want to put this on-device even if the hardware is there, unless there's some way of locking it down. I suspect we…
Non-tech enterprise was already doing this years ago.
Regulatory reasons, privacy reasons, security, etc. They want on-prem and total ownership of the data. Sometimes air-gapped.
Re: If this is true, the hyperscalers are toast
#80Earlier quoted context omitted.
Yes, it's all about the amount of money invested right now. Justifying that was always going to be tricky, and it still is. The compute capacity, however, is here to stay. There will be a seemingly endless line of people queued up to buy it at pennies on the dollar if this whole thing blows up financially. And they will use it!
Indeed, "pennies on the dollar" is the key point the author is making in this article.