Live data from Hacker News

Local AI needs to be the norm

unix.foo

691–700 of 804 posts

Re: Local AI needs to be the norm

#691

Earlier quoted context omitted.

> higher param count models will remain smarter for a looong time They're not smarter, they just know more stuff. You probably don't need knowledge about Pokemon or the Diamond Sutra in your enterprise coding LLM. The "smarts" comes from post-training, especially around tool use.

If the smarts came from post-training, we could show significant gains by doing that post-training again for previous generations of models. But we know that isn’t happening - effective post training is necessary but not sufficient for model performance.

> we could show significant gains by doing that post-training again for previous generations of models

That's what Chinese models are doing, and beating Opus et al.

Re: Local AI needs to be the norm

#692
post #408

People want local AI, but only if UX is good. Tooling/harness quality may matter as much as model quality. I think the future will probably be a hybrid of: 1. local AI for simple, private, everyday tasks 2. online AI for very hard or long tasks

The Clippy app someone made and posted here a while back is the perfect average person LLM interface; https://felixrieseberg.github.io/clippy/

This is so good. Wiring in small models for a variety of tasks would make this absolutely sing.

Re: Local AI needs to be the norm

#693

Earlier quoted context omitted.

No, exactly the opposite actually. Qwen3.6 is too imprecise for long running agentic tasks. It doesn't have the same ability to check itself as Gemma does in my testing. I keep Qwen MoE in vram by default because there are tons of tasks i trust it to oneshot and it's 90tok/sec is unparalleled, anything where I don't want to have to intervene too much it can't be trusted.

Oh interesting. I've read that Gemma 4 is really good for creative stuff, but I'm mostly interested in agentic coding. Unfortunately, each time I use Gemma 4, I just get it stuck in loops.

This is probably a precision thing, I think there's a really big difference in long running tasks between q4 and q6.

Re: Local AI needs to be the norm

#694

Earlier quoted context omitted.

In my experience once you get to ~30 gigs of ram for a model like Gemma4, the rest of the 128g of memory is simply nice to have. The speed and costs are what make it tough though, because its slower and more expensive than the same model served on a big accelerator card, and is going to be worse than a frontier model.

I wonder if it really needs to be worse. I am playing with the idea of fine tuning a model on my exact stack and coding patterns. I suspect I could get better performance by training “taste” into a model rather than breadth.

That approach has its advantages, but sometimes I want to generate code for a language or kind of project I’m not experienced with using the accepted best practices.

Re: Local AI needs to be the norm

#695

Earlier quoted context omitted.

Outside of the USA it would not look like a wealth transfer to an oligarch. Not every country is in a crypto-libertarian race to hoard power and wealth.

Not every country is in a crypto-libertarian race to hoard power and wealth. Meanwhile, in the EU, the model would be collectively financed, trained by a competent, neutral agency... and then completely lobotomized in the name of "the children," "safety," "IP rights," "correct speech," dozens of individual countries' legal and regulatory requirements, and any number of additional vocal, noncontributing NGOs. So no on…

European models are competitive, despite the concerns you raise.

I don't need a model that can easily produce CSAM or reproduce copyrighted works verbatim in order to be productive.

Re: Local AI needs to be the norm

#696

Earlier quoted context omitted.

False. The absolute capability is irrelevant, with the proper harness 31b is more than adequate for a very large portion of the tasks I ask AI to do. The metric isn't how good the model is at Erdos Problems, it's how reliably it can remove drudgery in my life. It just autonomously reverse engineered a bluetooth protocol with minimal intervention, it's ability to react to data and ground itself is constantly impressiv…

Maybe reaching for an analogy would be helpful here. Thot_experiment is saying that his 2016 Toyota Prius is a great and reliable car for his daily commute and running errands. Whereas everyone is screeching about its capability gap with a Lockheed Martin F35 lightning.

Yeah, thanks, though I think local models are at least a Cessna, which while being nothing like an F-35 can fly.

Re: Local AI needs to be the norm

#697

Earlier quoted context omitted.

Has anyone tried to calculate the break even cost of buying a PC to run an LLM locally, vs the amount of tokens you could get from an AI provider?

The basic answer: very much not worth it at face value, becomes arguably worth it once you start worrying about future rug pulls from the big AI providers. (And that does include the market for third-party inference, at least at present.) It's also worth it if you have existing hardware to repurpose, but that's obvious and not what you were asking about.

Also you can feed it ALL of your data willy nilly without ever worrying about safety because you can just do it with the LAN cable unplugged, for applications that demand data hygiene it's a cheat code that guarantees safety without any sort of data sanitization.

Re: Local AI needs to be the norm

#698
post #469

Cool, well let me know when Opus 4.5 level performance is available locally, at speeds that serve everyday use, and 100% I'm right there with you. Until then, I'm going to keep sending my JSON to the server farm in Virginia because it's the only place that can serve me a model that actually works for my uses.

Local models embody the hacker spirit, constant Claude glazing is spiritually incompatible with tinkering. Don't upload your spirit to the cloud.

Yo, MTP for Qwen is sick, thank you! Your work is invaluable.

Re: Local AI needs to be the norm

#699

Earlier quoted context omitted.

Have you actually tried this stuff or are you just saying stuff you hear on the internet?

Yes. I have tried this stuff. I really don't see how my use, or non-use of AI APIs changes this reality. Github Copilot announced it's going to per-token pricing in less than a month. I heard that on the internet by the way.

I'm watching people host models on things like LangSmith and OpenRouter for a fraction of the cost you are talking about. We have other people reporting their M4 Macs providing them with performance close to what they get with ChatGPT and Claude all locally with just a 24 GB M4 Mac. We already spend money on laptops. I can put in a ticket for an M4 Macbook Pro from IT right now.

Re: Local AI needs to be the norm

#700

I'm betting my startup on it. The subsidised model subscription will start to dry out and providers will lean heavier into locking down how they want their models to be used (Anhropic has been paving the way already). The only way forward is open weight models. If you are working on any LLM powered product be careful betting on utilising user subscriptions.

Maybe you know something I don't, but it seems the standard will continue to be a large number of companies hosting and reselling LLMs as both subscription plans and pay-as-you-go. It's virtually identical to the mobile market: the economics of the business require a large regular infusion of cash, and limits are used to prevent a minority of users from making the service unusable/unprofitable. A few giants are the most expensive but offer the most features, and cheap providers offer less for less. All of this will happen because people constantly want "more": more bandwidth, more quality, etc. Capitalism rewards this constant growth/advancement with constantly increasing bills.

Anthropic is going to go out of business by probably Q1 2027 due to not paying their bills. OpenAI will become a new Oracle, serving a luxury product for enterprises and governments. Google and Microsoft will keep doing what Google and Microsoft do. Chinese vendors will capture a significant amount of business over the next 10 years by running the models in non-Chinese DCs, with demand coming from their much lower prices. 95% of regular users will be paying for open model subscriptions, even if their local machine can run the model, because the providers will be offering features that are hard to impossible to replicate locally.

Post reply on HN