Live data from Hacker News

Local AI needs to be the norm

unix.foo

721–730 of 804 posts

Re: Local AI needs to be the norm

#721

Earlier quoted context omitted.

Both my experience, and Anthropic's off-peak promotion, indicate that there are very uneven levels of demand for peak hours versus off-peak hours. How close do you think they are?

But that's demand for cloud inference that's priced on a flat-rate basis with some adjustments (like "off-peak hours"). Not a local rig where inference is effectively free aside from the cost of power whenever the system isn't congested.

The local rig is not free and requires very large capital expenditures while producing very low token throughput for large models. Within any time budget, you can get many orders of magnitude more large-model tokens off an 8xB200 than off a local rig. Therefore cloud tokens have a huge capital efficiency advantage over local rigs. That will continue basically forever, since there will always be large cloud companies willing to spend millions of dollars for more capital-efficient hardware, so Nvidia and friends will continue to spare no expense producing it, meaning the cloud hardware will be way too expensive if you're not a large inference company. You can also buy local rigs, but they will be less capital efficient per token, not more.

(This is a generous argument: it also ignores the massive software stack optimization the cloud companies do that doesn't trickle down to local-rig-sized deployments; for example, prefill/decode disaggregation, which would double the VRAM requirements for a local rig — if you could even do it on a local rig, which you can't, because local rigs don't have Infiniband. But at scale, prefill/decode disaggregation improves capital efficiency, since you can tune the compute-bound prefill node differently than the memory-bound decode node.)

The advantage of local rigs is not capital-efficient tokens. It's privacy. But then again, you can get zero-data-retention options from many inference companies, so for many use cases it may not matter unless you need strict guarantees the data never leaves the building...

Re: Local AI needs to be the norm

#722

How is having local AI going to produce a result that's any better than using OpenAI or Anthropic? Isn't what we really need programmers who rely on themselves more than AI so they avoid technical debt accumulation?

Having local AI as a credible threat will keep them on their toes. Which will benefit consumers a lot.

Re: Local AI needs to be the norm

#723
post #311

Earlier quoted context omitted.

Jokes on you. We are already running Deepseekv4Flash, Mimo2.5, MiniMax2.7, Qwen3-397B locally in very affordable hardware. These models are in the real of Opus4.6. For those of us a bit crazy, we are running KimiK2.6, GLM5.1 and more ...

They all still fall short of Opus 4.6, definitely though. They are good but fail on extremely complex tasks, in contrast with a frontier model that will keep on trying until it succeeds or exhausts the solutions space.

frontier models don't keep trying until they succeed. that's a harness problem and best believe it, the best harness are private and not public.

Re: Local AI needs to be the norm

#724

Earlier quoted context omitted.

Two Mac Studio M3 Ultra 512GB and 1 USB cable can run all those models - maybe about $30,000 in hardware - and based on my benchmarks, those Mac Studios were twice as fast as the A100s on Deepseek v4 Flash, which has a quantization but not really a lossy one.

That cannot run KimiK2.6 or GLM5.1 i.e models within the ballpark of anything offered by frontier companies.

I run kimik2.6 and GLM5.1 on less than $10,000 system. Granted I started putting my system together 2 years ago when things were much cheaper. I run DeepseekV4Flash with 1 million context locally.

Re: Local AI needs to be the norm

#725

Earlier quoted context omitted.

These astronomical AI data centers will be used for high-value inference with smarter models that really are too large for running locally. The investments will be fine once they pivot to that use. Currently available open models are not in that range.

I don't buy that that will be a useful distinction. First of all, no AI model will say "I'm too smart for this question, I suggest you use a cheaper one so I don't make unnecessary money for my owner" or "I'm too dumb, so instead of hallucinating I'll suggest you go to the cloud and ask my smarter sibling". Second, there is no incentive in the market for tooling to evolve that way. There will be the illusion that som…

Dynamic routing is the usual name for the piece that orchestrates which LLM will be used, based on query complexity. There in an open source implementation as part of the vLLM project (and probably others), it is a field of active research in several universities and labs. It is also suspected that the frontier LLM providers might already be doing something like it behind the covers.

Re: Local AI needs to be the norm

#726
post #573

Earlier quoted context omitted.

I'm sorry to spoil it for you, but Perl script was able to do all of that like ... 10 years ago? The out-of-the-box Shotwell manages photos quite well without any intelligence. The problem, as people mentioned above, is SOTA models cognitive and tooling abilities. Also, have you noticed as top-end Mac Studios got downgraded recently? They don't want you to have access to frontier models. And you will not have it. See…

> They don't want you to have access to frontier models. And you will not have it. See Mythos as Exibit A. "They" fully well know that they current frontier model are maybe 6 month ahead of what people will have access to without their control. See Deepseek as Exibit B The reason you can't run these locally are more with the fact that those mythos sized models require extreme amount of memory and processing power to…

Isn't Mythos that screw up where Anthropic failed to ship something that was no better than the product OpenAI launched a few weeks later?

And, assuming the allegations are true, don't things like Deepseek and Qwen offer existence proofs that frontier models are (and will forever be) trivially distilled down to run domain-specific tasks on boxes that cost a few months of Claude Max subscription?

Re: Local AI needs to be the norm

#727
post #475

Earlier quoted context omitted.

Curious, why did Zed with ACP not work for you?

I'm just guessing, but IDE which is using 3D acceleration just for stupid UI to run "smoothly", that is ridiculous. Who runs IDE with LLM agents accessing your local filesystem, on bare metal? Or am I alone to run everything LLM related on my VM just for development work. Then because of ZED genius decision, you need to share your GPU to VM, then some important features will not work, like snapshots. So you also need…

What's wrong with using a 3d accelerator and falling back to CPU graphics if needed? Pixels / joule is orders of magnitude better on an iGPU than on the CPU. (Which can matter over a 8-12 hour editing session, maybe.)

Re: Local AI needs to be the norm

#728

Earlier quoted context omitted.

>It's here, right now. I mean I've been forcing my good old 1080ti to run local models since a short while after llama was first leaked. But I wouldn't say "local models are here" in the same way as "year of the Linux desktop!111" Until someone can just go out and buy some sort of "AI pod" that they can take home, plug in and hit one button on a mobile app to select a model (or even just hide models behind various pe…

What is the use case you see for non-technical users self-hosting? I think it’s important that tools remain available but I don’t expect it to be adopted by “average consumers.” I’m interested in self-hosting for privacy and control. I already owned the hardware I’m testing with, so my spend is limited to time and electricity. The “LLM pods” you describe will be loaded with spyware and adware (see: Smart TVs), and av…

And on top of that, I'm sure the "LLM pod" will still be sold on a subscription model so you get model updates etc.

But I wish we could actually have nice things. I imagine there's a niche for a middle ground: a privacy-preserving device that uses local-only models and doesn't spy on the user, and sells for a one-time payment with no subscription. It'll be expensive, though, likely more expensive than using a cloud-hosted model.

Re: Local AI needs to be the norm

#729
post #247

Earlier quoted context omitted.

Perhaps I am the odd one out here, but a small part of me wants to see what happens when you run a proprietary SOTA model on a laptop.

Currently I'm testing something like this just to see what happens. I have an old laptop with 4GB of RAM. I attached a USB drive with Gemma 4 31B model (which is 32.6 GB). Currently the laptop is running llama.cpp and trying to respond to a prompt by streaming the model from disk. The USB drive light is flickering, showing something is happening. It's been about 8 hours since I entered the prompt and I've gotten abou…

[deleted]

Re: Local AI needs to be the norm

#730
post #247
post #212

Earlier quoted context omitted.

> They will be, and that moment is not that far off. It's here, right now. I'm running quantized Qwen and Gemma on a decent, but three years old gaming rig (think RTX 3080 12GB and 32 GB RAM). Yes, it's slow, it has a small context window. But it can (given a proper harness) run through my trip photos and categorize them. It can OCR receipts and summarize spendings. It can answer simple questions, analyze code and ev…

Perhaps I am the odd one out here, but a small part of me wants to see what happens when you run a proprietary SOTA model on a laptop.

I don't think you're the odd one out. I would be very curious to try to run Opus 4.7 on a (high end) laptop. I'd also like to see how it runs on a high-end workstation rig built for it.
Post reply on HN