Live data from Hacker News

Running local models is good now

vickiboykis.com

601–610 of 651 posts

Re: Running local models is good now

#601

Local models Inference never really took with non tech people. Everyday people who dont know the difference between autocorrect and GPTs. But thanks to recent hardware launches from Nvidia and AMD's in response to the MacMini series, it is quite evident that local AI will replace the conventional laptop market completely one day. Current laptops will be what Nokia represents to a iPhone or Android user. Huge Leaps ah…

Are you referring to Spark that has 128 GB of Unified Memory? It would be still expensive.

Re: Running local models is good now

#602
post #136

I don't know about good, I use a lot of local models and they're still pretty painful to run locally You have dense models (qwen 27b, gemma 31b) who are pretty smart, but pretty slow You have MoE models (gemma 26b, qwen 35b, north mini code 30b) who are pretty fast, but make a lot of mistakes You need a lot of memory to run these well, quantization makes tool calling weaker, so most run at 4 bit quants and are wonder…

This is basically my experience as well. I have a moderately recent but high spec desktop (Radeon 6900 XT with 16 GB VRAM, Ryzen 9 7900X 12-core, 64 GB system RAM), and I tried out some recommended models with ollama a month or two ago. Anything not geared specifically towards coding seemed to struggled with actually making tool calls instead of just stating the actions they would take without making them (and trying…

> qwen refused to believe that it was running in ollama and insisted that it was running from the Alibaba cloud without access to my local system

Sounds like you were either running at a too-low quant, or you were trying to do Agents with something like Qwen 3.5 9B? Qwen 3.6 27B at Q4_K_M I can have that at running all night after a single one-shot a, Anne when I come back in the morning, it’s done

Re: Running local models is good now

#603
post #516
post #31

Earlier quoted context omitted.

Opus in my experience is equally unpleasant "character"-wise, but at least it actually gets stuff done more often, so it's at least slightly more earned at that. It's still a neurotic cargo-culting dogmatic idiot, but one that at least sometimes does produce deliverables instead of only bottom-tier HN-esque opinions. Hmm. I think I might just fundamentally disagree with Anthropic about the idea of what a "tool" shoul…

This morning I have been blessed with an example of the exact behavior that is so infuriating. > But re-reading the comment: > "In the real world however, it does not. Hence, in the future we might fail this check even if it works within this limited check." > The comment says "we might fail this check even if it works" — implying the original intent was to always fail (return 1) as a conservative stance, leaving roo…

Well there are in a business of selling you quote ubquote, intelligence, so it does what they want it to do: imply intelligence, which for me is just an illusion. But on the other hand this is what politicians and other conmen are doing? Using a lot of words and say nothing of value?

Also: more tokens.

Re: Running local models is good now

#604
post #15

After having been a happy user of Qwen3.6-27B for a few weeks, due to being away from the hardware, I'm currently forced to use Claude Sonnet 4.6 It is such a downgrade. I don't understand how that's even possible. The thing has so many strongly-held opinions I did not ever ask it for, talking just way too much and generally feeling somehow dumber. Of course, being significantly larger, it will encode more knowledge,…

Re being away from the HW: with Tailscale and llama-server it's now super easy to just run an inference server at home and use it from wherever you are.

Exactly what I do!

You can have web UI for PI!

Re: Running local models is good now

#605

Earlier quoted context omitted.

Not everyone has the right hardware.

I guess I’m thinking of the $100/mo users, for whom it’s probably possible to get the right hardware.

The hardware to match Opus costs at least $200k, and you have to maintain it, and it's still not going to be as good.

Re: Running local models is good now

#606
post #15

After having been a happy user of Qwen3.6-27B for a few weeks, due to being away from the hardware, I'm currently forced to use Claude Sonnet 4.6 It is such a downgrade. I don't understand how that's even possible. The thing has so many strongly-held opinions I did not ever ask it for, talking just way too much and generally feeling somehow dumber. Of course, being significantly larger, it will encode more knowledge,…

How qwen3.6:27b compare to qwen3.6:35b-a3b (MoE) in your experience (if you tried). I find the dense models are way too slow on my H/W.

[dead]

Re: Running local models is good now

#607
post #88
post #15

After having been a happy user of Qwen3.6-27B for a few weeks, due to being away from the hardware, I'm currently forced to use Claude Sonnet 4.6 It is such a downgrade. I don't understand how that's even possible. The thing has so many strongly-held opinions I did not ever ask it for, talking just way too much and generally feeling somehow dumber. Of course, being significantly larger, it will encode more knowledge,…

what kind of hardware do you need in order to run qwen3.6-27b

I bought r9700 for about 1700-1800$ and I have like 800t/s prompt and about 50t/s of inference on average? It hurt a bit when you change a prompt so llama.cpp have to discard entire cache and it have to think for 2-5min depending on the context, but otherwise it is faster than I can read.

Re: Running local models is good now

#608
post #82

Earlier quoted context omitted.

The opposite of that has been happening for 20 years now with cloud compute. It won't happen with AI models either. It's almost ingrained in the American business model now. Outsource everything. Nobody wants to manage a room full of servers when they can spend 2-3x as much and outsource that headache along with the responsibility for it. Same will happen with AI. Whether that means paying Anthropic that premium or p…

> The opposite of that has been happening for 20 years now with cloud compute. It won't happen with AI models either. AI is different. Cloud computing genuinely is cheaper on average. It's better than paying for cisco servers, and at scale, it's cheaper than managed platforms (ala Heroku), and it's a coin toss for when you're in the middle ground and constantly approaching the point of rebuilding poor-man versions of…

AI is different because you can't encrypt it. An running on someone else's hardware is basically just 'trust me bro! I won't read it!'. Of course you can say that about I.e. database too, but at least you can run it on your own dedicated hardware in some datacenter, so it is password protected, you can encrypt it at rest and you will only know the key.

With AI, no, you can't . model needs plain text to be able to work. If somebody will be able to figure out models with asymmetric keys will make a lot of money.

Re: Running local models is good now

#609
post #549

Earlier quoted context omitted.

Then I'm interested if there are any facts as to what ZDR actually means?

It can still mean Zero Data Retention - i just comes down to whether you trust the company to actually do what they promise. The fact that they've trained models on data that wasn't theirs does not make me trust them a lot when they make this claim.

Once they feed your data into the training dataset, they can delete the individualized copy. The training dataset is, of course, a trade secret that can never be exposed without causing serious harm to the company's model, or equivalent legalese that will prevent it's disclosure to all, governments included.

Re: Running local models is good now

#610

Earlier quoted context omitted.

When discussing this, may I ask (I know you are probably bored of the actual arguments), what does "trained models on data that wasn't theirs" actually mean in practice? Again, I know these arguments have been done to death, but every human who reads source code that wasn't written by them, or views art that wasn't created by them, and practices against this art, is training their brain on data "that wasn't theirs".…

A product is not a human. They are selling a product based off copy-righted material without the rights to it. It's a pretty easy line to draw, honestly.

Why is something being human or not relevant here?
Post reply on HN