Live data from Hacker News

Local AI needs to be the norm

unix.foo

491–500 of 804 posts

Re: Local AI needs to be the norm

#491

I do think local models are the future, but there's still the question of cost to be answered. Even if there's some slew of effincency improvements that mean an LLM can run locally on consumer level hardware on an affordable budget (and that's a big "if"), there's still the cost of training the modles to consider. Assuming we end up in a future where people pay to run multiple smaller models on their machines for spe…

> the question of cost to be answered.

Commoditizing complements. If Anthropic/OpenAI/etc is eating your lunch, make it work with cheap local LLMs , you can beat them on price by having local inference you don't pay (nor need data centers for), and try to keep your (user/data) moat.

The more Anth/OAI disrupt, the more likely this is to happen. If they don't disrupt enough (.ie: grow as an ecosystem to defend against incentives to commoditize), then yes, those incentives are removed, but they also leave money on the table, which they need.

Not only at business level, but also geopolitical (to a lesser extent? or not since lots of open weight models comes form China?).

Re: Local AI needs to be the norm

#492
post #212
post #80

They will be, and that moment is not that far off. We've got the progression in place already: first, large data centers could have performant LLMs, we are now firmly in "a bunch of servers with a couple of H100s each" territory, slowly going into "128 GB VRAM on a MacBook Pro or a Strix Halo". Within the next year, the pattern of "expensive remote LLM for planning, local slow-but-faster-than-human LLM for execution"…

> They will be, and that moment is not that far off. It's here, right now. I'm running quantized Qwen and Gemma on a decent, but three years old gaming rig (think RTX 3080 12GB and 32 GB RAM). Yes, it's slow, it has a small context window. But it can (given a proper harness) run through my trip photos and categorize them. It can OCR receipts and summarize spendings. It can answer simple questions, analyze code and ev…

I'm sorry to spoil it for you, but Perl script was able to do all of that like ... 10 years ago? The out-of-the-box Shotwell manages photos quite well without any intelligence. The problem, as people mentioned above, is SOTA models cognitive and tooling abilities. Also, have you noticed as top-end Mac Studios got downgraded recently? They don't want you to have access to frontier models. And you will not have it. See Mythos as Exibit A.

Re: Local AI needs to be the norm

#493
post #487

Earlier quoted context omitted.

Flat wrong. Q6 Gemma 31b feels a lot like opus 4.5 to me when run in a harness so it can retrieve information and ground itself. The gap is not that big for a lot of usecases. Qwen MoE is fast as fuck locally for things that are oneshottable. I have subscriptions to all the major providers right now and since Gemma 4 and Qwen 3.6 came out I haven't hit limits a single time. I'm actually super surprised by the number…

What harness are you using ? I'm going to switch to local LLMs for most stuff soon.

Overall using screentime as the metric, derived from some imperfect logging and vibes it's about 50% OpenCode 15% Continue 15% my homebrew bullshit 13% Claude Code and 7% Cline. I've been deep on agentic stuff lately (1.3wks aka 3 months of AI time), there are only so many hours in the day to duplicate work and AB test, but in the past I've sworn by Qwen Coder + llama.vim and I still enjoy that workflow for deep work far more than I like prompting agents, but there's a lot of dross I'm learning to delegate.

Re: Local AI needs to be the norm

#494

Earlier quoted context omitted.

I've never had a "What's new" tab ever open because I disable the customized home page where that's displayed. I'm guessing you're not aware that's an option. Please show me where in either of those documents it explains it's going to download a 4GB model.

I use an extension that gives me a customized homepage, but I still always get the "what's new" tab on every major version upgrade. It's a totally separate tab that opens. It's got nothing to do with what you use as your homepage.

Thank you for going out of your way to deny my exact experience. Do you think I'm doing this to rag on Google? And you're this eager to defend them?

I'm on gentoo. I have to update chrome manually. I updated it. On update I _never_ get a "what's new" page. I've had this profile for more than a decade so I have no actual idea why, but, I can absolutely tell you, I do *not* get one. After update it started consuming all my bandwidth. This use did not show in it's task manager. I have a metered connection. This is a problem for me. I worried it was a compromised plugin. I had to spend 10 minutes in Firefox discovering why chrome was doing this then going to the configuration and disabling this.

This was a disappointing experience. I'm sorry you feel differently; other than stating the obvious, I seriously have no idea what you and the other corporate defense squad members are trying to achieve with this gaslighting nonsense.

Re: Local AI needs to be the norm

#495

Opus 1M context window and lighting fast response time is hard to compete with, even if you run a local A100 the local models are just not as good as tool calling, long running tasks and non-hallucinations

It was hard for an Apple ][ to compete with an IBM mainframe at enterprise data processing, but the power of personal ownership & commodity economics was disruptive enough that 30 years later 99%+ of enterprise data processing was taking place on descendants of the original personal computers.

Re: Local AI needs to be the norm

#496
post #25

Local models are extraordinarily expensive if you're not maximizing throughput, and you're not going to be maximizing it. Local models need to be resident in expensive RAM, the kind that has fat pipes to compute. And if you have a local app, how do you take a dependency on whatever random model is installed? Does it support your tool calling complexity? Does it have multimodal input? Does it support system messages i…

I don't know why you are being downloaded. These are precisely the facts that advocates for local models completely ignore. Local models are absolutely going to be the future for things like simple automation and classification tasks that run occasionally and don't need to rely on internet access. But for all of the serious stuff where you are doing knowledge work, the models will simply continue to be too big, and t…

Downvotes: IKR? It's a signal that there's a lot of motivated reasoning.

I don't think that many people have built apps against these models.

I mean, I use a heavily quantized version of qwen3 for image classification, caption generation, prompt expansion etc. for image generation, instruction-driven edits, and so on. You can go a long way when you don't need a lot.

A model that can do tool calls - any tool calls at all - can look reasonably cool once you put it in a harness where there's enough immediate context to take action. You can get carried away by anything happening at all. But golly gosh it's a long way short of intelligence available in the bigger models.

And the lighter you make your harness, giving the model more free reign, more autonomy, you get a big jump in capability combined with a big jump in failure modes when the model is dumb.

Re: Local AI needs to be the norm

#497

Earlier quoted context omitted.

API prices are most likely not subsidised. A brief look at openrouter can tell you that. There are plenty of providers that have 0 reason to subsidise that sell models at roughly the same average price. So the model works for them (or they wouldn't do it otherwise).

They are subsidized, heavily. This is simple math, there are lots of reasons to subsidize. Please go look up the hardware requirements to run your favorite model and a given tok/ps then multiple that by 86400 (seconds in a day) then divide that by 1mm and multiple by the $ per mm tokens, then ask yourself if there's any possibility they could be profitable or even close to break even. You are going off vibes alone, t…

Serving a single user is likely not profitable, but total throughput rises a lot when serving many concurrent users, because the same weights can be used to generate tokens for all users at once, which increases efficiency.

Also, a lot of money is being made on input tokens and cached tokens, which are much cheaper to compute.

DeepSeek published their math for serving the V3/R1 models. They were 535% profitable: https://github.com/deepseek-ai/open-infra-index/blob/main/20...

Re: Local AI needs to be the norm

#498

I do think local models are the future, but there's still the question of cost to be answered. Even if there's some slew of effincency improvements that mean an LLM can run locally on consumer level hardware on an affordable budget (and that's a big "if"), there's still the cost of training the modles to consider. Assuming we end up in a future where people pay to run multiple smaller models on their machines for spe…

> the question of cost to be answered. Commoditizing complements. If Anthropic/OpenAI/etc is eating your lunch, make it work with cheap local LLMs , you can beat them on price by having local inference you don't pay (nor need data centers for), and try to keep your (user/data) moat. The more Anth/OAI disrupt, the more likely this is to happen. If they don't disrupt enough (.ie: grow as an ecosystem to defend against…

What are you talking about Willis?

Re: Local AI needs to be the norm

#499
post #408

People want local AI, but only if UX is good. Tooling/harness quality may matter as much as model quality. I think the future will probably be a hybrid of: 1. local AI for simple, private, everyday tasks 2. online AI for very hard or long tasks

it's a self enforcing loop

local LLMs builds tool that does exactly what user wants, how it wants it, which is bext UX

this becomes AI literacy

LLMs already nicely bridge the gap form "I want this" to "here's a local page that does it".

examples of tools i have built that requires almost very low tech knowledge * push a button on my phone to take screenshot in my mac (when i watch videos) * help me exercise, gamify it for me * "help me track time spent online to how it impacts what i do in real life, built a tool that rewards and me points me towads things that make me DO things online" * i want to improve my writing, give me exercises and build addiitonal tools (leading to an "append only" digital keyboard i use to exercise )

local AI can already create these tools, and no external company is ever going to beat me/the-user because instead of getting features i don't want, or that almost do what i want, or that do something that advantages the company they just do what I want

Repositories of tools-as-ideas created by others are quite often just index.html and ... that's all? manage data in localstorage, end of it?

Online inferences is still needed for large data (audio/video/images) processing. For now? we don't know, history suggests we'll have the capabilities to do that locally "soon". Or maybe not :)

The main issue is "online for collaboration". Not same user across different devices, that is easy. MeteorJS-style approaches (making local copies of part of dbs, reconcile to remote/origin) seems to be an interesting possibility at small scale, since once you have the right primitives in place you can go horizontally everywhere.

Re: Local AI needs to be the norm

#500
Agree with the sentiment, but: "We are building applications that stop working the moment the server crashes or a credit card expires."

This has been the case for way longer than openAI and Anthropic has been around with services like AWS, Cloudflare, etc.

Post reply on HN