Live data from Hacker News

Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

qwen.ai

391–400 of 482 posts

Re: Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

#391

I'm kind of interested in a setup where one buys local hardware specifically to run a crap ton of small-to-medium LLM locally 24/7 at high throughput. These models might now be smart enough to make all kinds of autonomous agent workflows viable at a cheap price, with a good queue prioritization system for queries to fully utilize the hardware.

I would love to have a shit load of small (27B dense. 35B MoE) agents running locally and looking at and ingesting every bit of data about me, my life and what I get up to see what sort of correlations it finds. Give a coding agent access to a data lake of events and let it build up its own analytics tooling to extract and draw out information from that data, and present it to me as daily/weekly/monthly summaries.

[deleted]

Re: Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

#392
post #15

Unsloth quants available: https://unsloth.ai/docs/models/qwen3.6

Getting ~36-33 tok/s (see the "S_TG t/s" column) on a 24GB Radeon RX 7900 XTX using llama.cpp's Vulkan backend: $ llama-server --version version: 8851 (e365e658f) $ llama-batched-bench -hf unsloth/Qwen3.6-27B-GGUF:IQ4_XS -npp 1000,2000,4000,8000,16000,32000 -ntg 128 -npl 1 -c 34000 | PP | TG | B | N_KV | T_PP s | S_PP t/s | T_TG s | S_TG t/s | T s | S t/s | |-------|--------|------|--------|----------|----------|----…

Did you try GPU/CPU mix with a bigger model?

Re: Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

#393
post #133

Earlier quoted context omitted.

question: why not use something like Claude? is it for security reasons?

We do make Claude and Mistral available to our developers too. But, like you said, security. I, personally, do not understand how people in tech, put any amount of trust in businesses that are working in such a cutthroat and corrupt environment. But developers want to try new things and it is better to set up reasonable guardrails for when they want to use these thing by setting up a internal gateway and a set of rea…

Because it's a great tool and the second it's not we can just do what you're doing :)

Re: Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

#394
post #106

The pelican is excellent for a 16.8GB quantized local model: https://simonwillison.net/2026/Apr/22/qwen36-27b/ I ran it on an M5 Pro with 128GB of RAM, but it only needs ~20GB of that. I expect it will run OK on a 32GB machine. Performance numbers: Reading: 20 tokens, 0.4s, 54.32 tokens/s Generation: 4,444 tokens, 2min 53s, 25.57 tokens/s I like it better than the pelican I got from Opus 4.7 the other day: https://si…

PelicanBench, the last benchmark for AGI.

Re: Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

#395

Earlier quoted context omitted.

> For coding often quality at the margin is crucial even at a premium For some problems, sure, and when you are stuck, throwing tokens at Opus is worthwhile. On the other hand, a $10/month minimax 2.7 coding subscription that literally never runs out of tokens will happily perform most day-to-day coding tasks

"Literally never runs out of tokens?" lol, no. Tokens are just energy. There is always a way to run out of tokens, and no one will subsidize free tokens forever.

if you run it at home then the sun is a pretty good way to get "free energy."

Re: Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

#396
post #15

Unsloth quants available: https://unsloth.ai/docs/models/qwen3.6

llama-batched-bench -hf ggml-org/Qwen3.6-27B-GGUF -npp 512,1024,2048,4096,8192,16384,32768 -ntg 128 -npl 1 -c 36000 M2 Ultra, Q8_0 | PP | TG | B | N_KV | T_PP s | S_PP t/s | T_TG s | S_TG t/s | T s | S t/s | |-------|--------|------|--------|----------|----------|----------|----------|----------|----------| | 512 | 128 | 1 | 640 | 1.307 | 391.69 | 6.209 | 20.61 | 7.516 | 85.15 | | 1024 | 128 | 1 | 1152 | 2.534 | 404.…

[deleted]

Re: Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

#398

Earlier quoted context omitted.

> For coding often quality at the margin is crucial even at a premium For some problems, sure, and when you are stuck, throwing tokens at Opus is worthwhile. On the other hand, a $10/month minimax 2.7 coding subscription that literally never runs out of tokens will happily perform most day-to-day coding tasks

"Literally never runs out of tokens?" lol, no. Tokens are just energy. There is always a way to run out of tokens, and no one will subsidize free tokens forever.

"Never runs out of tokens" in the sense that running 8 hours a day 7 days a week is still under the subscription limit

Re: Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

#399

Earlier quoted context omitted.

The 27B model they release directly would require significant hardware to run natively at 16-bit: A Mac or Strix Halo 128GB system, multiple high memory consumer GPUs, or an RTX 6000 workstation card. This is why they don’t advertise which consumer hardware it can run on: Their direct release that delivers these results cannot fit on your average consumer system. Most consumers don’t run the model they release direct…

I can run Qwen3.5-27B-Q4_K_M on my weird PC with 32 GB of system memory and 6 GB of VRAM. It's just a bit slow, is all. I get around 1.7 tokens per second. IMO, everyone in this space is too impatient. (Intel Core i7 4790K @ 4 Ghz, nVidia GTX Titan Black, 32 GB 2400 MHz DDR3 memory) Edit: Just tested the new Qwen3.6-27B-Q5_K_M. Got 1.4 tokens per second on "Create an SVG of a pellican riding a bicycle." https://gist.…

Don't forget that you're also spending much more electricity because it takes so long to run inference.

Re: Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

#400

Earlier quoted context omitted.

Why would you though? And by the way: Thanks for relentlessly holding new models’ feet to the pelican SVG fire.

Because I want to read about Qwen, not someone's one-off vibe test followed by 1:1 conversations. (case in miniature here: which is the last comment in this thread that says something about Qwen? The root post. Is that fun policing? Yes, apologies.)

I understand your reasoning and it's valid, but I think the best you can do is indeed collapse the thread (not sure if any mobile clients do better than that?)

It's perhaps not a serious test, it isn't to me, but on the edges of jokes about pelicans they're usually some useful things people smarter than me say, and additionally if providers are spending some time on making pelicans or svg look better, this benefits all of us.

So, no hard feelings, you're understood (and I'm not trying to be patronising, I'm just awkward with the language), but pelicans are here to stay because it seems that the consensus is they're beneficial and on topic.

All the best!

Post reply on HN