Live data from Hacker News

Qwen 3.6 27B is the sweet spot for local development

quesma.com

261–270 of 809 posts

Re: Qwen 3.6 27B is the sweet spot for local development

#261

Earlier quoted context omitted.

> Being able to nail a zero-shot greenfield project is relatively easy even for a small model Not really germane to your comment but I hope I don’t sound old when I say I remember a time when spinning up a PoC was a week of work, and a statement like yours was pure science fiction.

Yeah, and we still do take a week for people that actually care. If I start prompting away the core of a new project I lose interest in the entire thing almost straight away. I hate it. The next day I could care less about it. In fact it just makes me lazy, like a fat person who drives everywhere. I love typing code and thinking for myself. Im going to continue to do that. I still dont know anyone who's shipped anyth…

I mean, have you looked for examples of things that people using local models to build and ship? Or are you just assuming it doesn’t happen?

I’ve used Qwen 3.6 27B for many things at work, and I’m regularly able use it for reasonably scoped tasks.

I’m not saying these models are perfect.

But you are complaining about people on the extreme, while at the same shouting from the opposite extreme.

Re: Qwen 3.6 27B is the sweet spot for local development

#262

Earlier quoted context omitted.

I am considering getting something like NVIDIA's RTX Spark when it comes out, though even that will be limited to 128GB.

It's out, I'm daily driving one. It's great

I assume you have the dgx spark? At this point I am not 100% on the difference other than Linux and Windows. The RTX spark should come around Q4, unless I am mistaken.

Re: Qwen 3.6 27B is the sweet spot for local development

#263

Running 27B dense model on M5 128GB is ok, but one can do better. On M5 128GB one can make use of the ram and use sparse MoE. For example, DeepSeek-V4-Flash will fit, served by DwarfStar ( https://github.com/antirez/ds4 ). One will probably improve 2x the token/sec speed, given DS4F 13B activated params in the MoE are ~1/2 of the ~27B of the dense Qwen. 27B Of the Qwen fit even on a cheaper 24GB card, e.g. amd 7900xt…

Works beautifully on a 3090, very usable speed. Don't expect Opus 4.8-level performance, but there are some things you just need to keep local.

Re: Qwen 3.6 27B is the sweet spot for local development

#264

Earlier quoted context omitted.

If you want to do coding with a local LLM your best bet is a 6 year old Nvidia 3090 which is substantially more powerful than the highest end overhyped Apple product for 1/5th the price.

That’s 24GB VRAM. Not enough to run a 27B model at a useful quant+context size.

Yeah seems to me like the mac studios with the unified memory architecture are genuinely good bang for the buck at the moment, because of this memory size consideration?

Re: Qwen 3.6 27B is the sweet spot for local development

#265

Earlier quoted context omitted.

That’s 24GB VRAM. Not enough to run a 27B model at a useful quant+context size.

You can run 8bit 27B models at 24GB, it's definitely enough for the model size.

I think that’s only true for MoE models. A dense model like 3.6 27b will require more (plus a KV store).

Re: Qwen 3.6 27B is the sweet spot for local development

#266
post #212

I love my MacBook Pro M5 128GB RAM and I love qwen3.6. BUT DO NOT buy this MacBook if you plan on doing serious coding using local LLMs with it. The reason is simple: your fingers will burn and your head will explode from the noise. Running any kind of sophisticated job on the very laptop you are using is just not viable. Sure you can use it in clamshell mode, but forget touching it while working with AI coding or ag…

Would the new upcoming AMD AI ryzen halo desktop be a better value offer? or dgx spark? You would have to get a third party reseller/scalper or refurbished mac mini to get 64gb of ram ever since apple stopped selling it.

Check the LLM benchmarks once it's out: it's such a common use case for these kinds of machines, you won't be waiting long.

Re: Qwen 3.6 27B is the sweet spot for local development

#267
post #50

The article is based on running Qwen 3.6 on a 128GB MacBook Pro. For reference, a 128GB MBP currently starts at $6699 USD [0] Some people will be happy to pay that premium for privacy, but at roughly 10X the cost of a MacBook Neo, that money could also buy a lot of credits on OpenRouter or frontier labs. [0]: https://www.apple.com/shop/buy-mac/macbook-pro/14-inch-space...

The maths there is pretty undeniable, but it is not where I'd make the split. Having a machine that can run some modest local LLMs, like the Gemma 4 12B, is really worth it. I don't know how much serious hands-free agentic coding I will ever do on my MacBook alone, but I do know that I would not have got so far into understanding this without tinkering with local models, llama.cpp, LM Studio, and LM Studio and all th…

> Having a machine that can run some modest local LLMs, like the Gemma 4 12B, is really worth it.

Agree having a powerful machine is really worth it in general for professionals, but strong disagree that running local LLMs has anything to do with it. It's hard enough as it is getting a good ROI on your time/money prompting/wrangling with frontier models. IMO leaning on the comparatively limited capabilities of local LLMs is best avoided in favor of keeping your own personal coding skills fresh and continuing to learn new ones.

Re: Qwen 3.6 27B is the sweet spot for local development

#268

Earlier quoted context omitted.

The model they reference can be easily run with 24gb+ of VRAM, and there are other similar models capable of running easily on 16gb of VRAM. It's not like 128gb is a requirement here.

For a MBP I have 48 GB of RAM M5 Pro. It runs at about 12-14 t/s at Q4, you could probably optimize it further. RAM is not a limitation but overall memory bandwidth. Q8 is slower. 35B A3B Qwen is quite speedy, but a little less accurate. With Qwen 3.6 27B dense I can squeeze a 9B parameter model and use that for fast analysis or code scanning while 27B is churning on a task in the background. It is tight, but totally…

I was doing some benchmarking last night on 2 3090s. The systems but old but I’m seeing 11tks 27b, 15tks 35b MoE.

The limited context is problematic. I’m not exactly sure what it’s got available but hermes was hit and miss on a prospecting job.

It does seem to be doing useful work but it’s not API call level quality

Re: Qwen 3.6 27B is the sweet spot for local development

#269
post #231

I love my MacBook Pro M5 128GB RAM and I love qwen3.6. BUT DO NOT buy this MacBook if you plan on doing serious coding using local LLMs with it. The reason is simple: your fingers will burn and your head will explode from the noise. Running any kind of sophisticated job on the very laptop you are using is just not viable. Sure you can use it in clamshell mode, but forget touching it while working with AI coding or ag…

I have an M4 Max and when I was trying out local LLM work with pi it has probably felt like the hottest I've ever felt any kind of Macbook be. I could feel the radiated heat off it even a few inches away. Honestly felt hotter than any Intel Macbook I've used. Because of that I stopped as I didn't want to harm my laptop in case I need to hold it for 10 years due to all the supply issues/price increases.

I tried to run it on a M4 Air for shits and giggles.

After about 1 minute the entire machine basically bricked and I had to hard reset :D

Post reply on HN