Earlier quoted context omitted.
> I tried running a smaller model locally, and it's not usable for me. If you have the hardware, a MacBook Pro for Qwen 3.6 35B A3B and Gemma 4 26B A4B for example, they are absolutely usable, both in terms of speed and quality. Anecdotally, I can use Qwen for day-to-day coding tasks in TS and Go, without hickups.
You let a hiccup slip through in your comment though.
If this is true, the hyperscalers are toast
81–90 of 108 posts
Re: If this is true, the hyperscalers are toast
#82Earlier quoted context omitted.
I very much don’t want to run it locally. I want the same one running somewhere else that I can interact with from all my devices. Look at something like Grok Bot. Nobody is going to run this locally. You can already self host almost anything, yet most people and businesses don’t.
> I very much don’t want to run it locally. I want the same one running somewhere else that I can interact with from all my devices. I run models on my desktop and access them from a phone app on the go. Wireguard tunnel. Responds fast and lets me kick off tasks or workflows via text or voice.
Maybe I could use codex or whatever to set that all up nowadays, but it was such a pain doing the initial setup and getting it all running and finding sources of media yadda yadda like a decade ago when I last tried lol
Re: If this is true, the hyperscalers are toast
#83Earlier quoted context omitted.
> I very much don’t want to run it locally. I want the same one running somewhere else that I can interact with from all my devices. I run models on my desktop and access them from a phone app on the go. Wireguard tunnel. Responds fast and lets me kick off tasks or workflows via text or voice.
Exactly the type of thing I’d pay someone not to deal with. It’s like media streaming vs piracy. Even though I could get a nice local setup with plex and have higher quality streams with free pirated br rips, I just don’t want to deal with it. Maybe I could use codex or whatever to set that all up nowadays, but it was such a pain doing the initial setup and getting it all running and finding sources of media yadda ya…
Very much a personality and/or lifestyle thing, but I've also been building infra for decades. I mean it was two container startups and a phone app download (10min). Not exactly difficult. If I didn't like it I certainly wouldn't be in tech!
> pay someone
I'd do that with something like 3 story roof work or foundation work with a contractor, but tech? If I want the outcome I will be happy with, often have to do it myself.
Re: If this is true, the hyperscalers are toast
#84Earlier quoted context omitted.
Exactly the type of thing I’d pay someone not to deal with. It’s like media streaming vs piracy. Even though I could get a nice local setup with plex and have higher quality streams with free pirated br rips, I just don’t want to deal with it. Maybe I could use codex or whatever to set that all up nowadays, but it was such a pain doing the initial setup and getting it all running and finding sources of media yadda ya…
> Exactly the type of thing I’d pay someone not to deal with. Very much a personality and/or lifestyle thing, but I've also been building infra for decades. I mean it was two container startups and a phone app download (10min). Not exactly difficult. If I didn't like it I certainly wouldn't be in tech! > pay someone I'd do that with something like 3 story roof work or foundation work with a contractor, but tech? If I…
I was into home automation and self hosting media and some other stuff for a while. It’s the kind of thing I’d only do again if i was equally or more interested in the process than the outcome.
Re: If this is true, the hyperscalers are toast
#85Earlier quoted context omitted.
The valuations of the hyperscalars won't sustain just being more efficient than something you can run locally. There's a market there, but it's for margin on a commodity. They're priced for oligopoly on unique, premium products.
I very much don’t want to run it locally. I want the same one running somewhere else that I can interact with from all my devices. Look at something like Grok Bot. Nobody is going to run this locally. You can already self host almost anything, yet most people and businesses don’t.
There will be some “rising star” businessperson who somehow got an amazing deal on land and energy and just likes to spend 1/2 their time in Shenzhen or something, but the datacenter is in Vietnam or Singapore or Thailand or wherever. They might even have a US branch, to make everyone feel better!
It will collapse the market, and everyone will realize that data centers are actually worth LESS than Toyota - they’ll look more like Tulips, suddenly.
In my opinion, of course.
Re: If this is true, the hyperscalers are toast
#86One of the big things to think about is whether local LLMs will be things companies want to deploy. If you think of for e.g. some proprietary piece of software that wants to embed an LLM they've fine tuned or trained, they will want to make back some of their research cost right. So they are not going to want to put this on-device even if the hardware is there, unless there's some way of locking it down. I suspect we…
> One of the big things to think about is whether local LLMs will be things companies want to deploy. Non-tech enterprise was already doing this years ago. Regulatory reasons, privacy reasons, security, etc. They want on-prem and total ownership of the data. Sometimes air-gapped.
This is in the UK where there's currently a big focus on data sovereignty in general, and I'm genuinely surprised by how little that's spilled over into demand for inference sovereignty (so far).
I still expect demand to grow substantially, but I've been saying that for the past couple of years and am beginning to wonder if there'll need to be some sort of trigger event before it happens (eg. the datacentre bubble bursting, or some sort of major scandal).
Re: If this is true, the hyperscalers are toast
#87Earlier quoted context omitted.
The valuations of the hyperscalars won't sustain just being more efficient than something you can run locally. There's a market there, but it's for margin on a commodity. They're priced for oligopoly on unique, premium products.
I very much don’t want to run it locally. I want the same one running somewhere else that I can interact with from all my devices. Look at something like Grok Bot. Nobody is going to run this locally. You can already self host almost anything, yet most people and businesses don’t.
I'm absolutely going to run this localy if I can. The only reason I'm not doing it is the fact I'm literally priced out.
Re: If this is true, the hyperscalers are toast
#88The paper underlying this blog post is fundamentally flawed because of benchmark ceilings. If we define only simple tasks like asking what is the capital of France, all models will converge to 100%, obviously. But as bigger models get more capable we want them to replace more and more complex tasks, in as short time as possible. Then of course there is the economics of it. Do people prefer to spend $5000 upfront to g…
Re: If this is true, the hyperscalers are toast
#89That design paradigm sounds familiar
Everyday I use smalll applications that do only one thing, some written in the 1970's
This submission got [flagged]. Later the "[flagged]" label was removed
Re: If this is true, the hyperscalers are toast
#90Earlier quoted context omitted.
Hyperscalers don't run computing at some multiple more efficient than on prem. The only way hyperscalers can compete is if they own the sand to cycles supply chain and ensure that raw compute is out priced in the market (ram,flash,compute). Ram and flash were an easy target because they are a commodity in name only.
> Hyperscalers don't run computing at some multiple more efficient than on prem. I'd disagree here. I see two avenues for an efficiency multiple, albeit a single-digit multiple: * Client aggregation allows a hyperscaler to average out demand spikes from uncorrelated clients, reducing the peak:average demand ratio and allowing better budgeting of compute. * Dynamic batching allows typical requests to run in batches of…
Go get a job a hyperscaler, they want to smoke what you are smoking.
I am not talking about diurnal cloud workloads, I am talking about the native efficiencies of hyperscalers vs on-prem. They have no magic and they all think they are going to make up their business overheads in exorbitant saas pricing.