Live data from Hacker News

Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

qwen.ai

461–470 of 482 posts

Re: Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

#461
post #430

Earlier quoted context omitted.

IMHO looks more like a stork, not a pelican. Look up any image of an actual pelican and check the ratio of legs to body. IMHO that's a weird mistake to make when asked for a "pelican". Have you considered asking a couple of artists on Fiverr or something to draw you a picture with the same prompt? I don't mean this as a gotcha, it's actual advice, you should probably get a sense of what a real human artist/designer (…

Honestly it never crossed my mind to waste some artist's time with this, but now that the joke "benchmark" has somehow reached orbital velocity maybe I should be thinking about it! I've run the prompt through dozens of dedicated image generation models so I've seen many versions of this that are better attempts than a text model spitting out SVG - here's gpt-image-2 as a recent example: https://chatgpt.com/share/69ea…

I believe that if you pay them for their time, it's not really "wasted", at least not nearly as "wasted" as when the next person would pay them to design some vapid advertisement.

In addition to that 1) it's for science and 2) maybe you owe it to yourself to have a really nice framed picture of a pelican riding a bicycle on the wall :D

About the dedicated image generation results, I still would have made the bicycle smaller, but it starts to depend on how motivated the artist is to make both the bike and pelican accurate. Which is fine, but if you want to have a benchmark, it's important to have at least one "known good" example, I think.

Re: Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

#462

Earlier quoted context omitted.

> Qwen 3.6:27b uses 29/32gb of vram What context size are you using for that? Btw, are you using flash attention in Ollama for this model? I think it's required for this model to operate ok.

I squeezed it into 24 GiB VRAM (since I have RX7900XTX): -- Q5_K_M Unsloth quantization on Linux llama.cpp -- context 81k, flash attention on, 8-bit K/V caches -- pp 625 t/s, tg 30 t/s

I have the same GPU and get very good results, even better than Gemma 4 26B A4B, using the following setup (Fedora 43 Silverblue, podman compose):

  services:
    llama:
      image: ghcr.io/ggml-org/llama.cpp:server-vulkan
      container_name: llama-qwen3.6-27b-dense
      ports:
        - 4201:8080
      volumes:
        - ./Qwen3.6-27B-Q4_K_M.gguf:/models/model.gguf:ro,z
        - ./mmproj-BF16.gguf:/models/mmproj.gguf:ro,z
      devices:
        - /dev/dri
      group_add:
        - video
      command: >
        -m /models/model.gguf
        --mmproj /models/mmproj.gguf
        --alias "Qwen3.6 27b Dense"
        -ngl 99
        -c 98304
        -b 2048
        --host 0.0.0.0
        --port 8080
        --parallel 2
        --kv-unified
        --ubatch-size 2048
        --flash-attn on
        -cb
        --jinja
        --no-webui
        -ctk q8_0
        -ctv q8_0
        --image-min-tokens 1024
        --temp 0.6
        --top-k 20
        --top-p 0.95
        --repeat-penalty 1
        --presence-penalty 1.5
        --reasoning auto
      restart: unless-stopped

Re: Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

#464

Earlier quoted context omitted.

Qwen 3.6 27B, and other dense models, as opposed to MoE models do NOT scale well. Like I said in my original post, for 27B usage specifically, I'd take a dGPU with 32GB of VRAM over Strix Halo. I also don't usually benchmark out to 200k, my typical depths are 0, 16k, 32k, 64k, 128k. That said, with Qwen 3.5 122B A10B, I am still getting 70 tok/s PP speed and 20 tok/s TG speed at 128k depth, and with Nemotron 3 Super…

Thanks this is very helpful for planning out localLLM buy. Sounds like we are still at least 1 generation out (DDR6 500-700GB/s memory) from getting to that magic ~25-30TG/s. Nemotron 3 Super architecture sounds promising.

Medusa Halo is on my wishlist, but I'm hearing late 2027 :(

M5 Ultra may be a better near-term option, expected in June. Supposedly ~1.2 TB/s unified memory, unsure of whether Apple will revive the 512 GB SKU or limit to 256 GB, but the new Neural Engine in every GPU core should help dramatically. These were always compute limted rather than bandwidth limited, even in M3 Ultra era.

The big cost of course being that you're locked into Apple silicon and Apple's walled garden. You can still use MacOS without creating an Apple account... for now...

At least Apple Silicon holds resale value remarkably well.

Re: Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

#465
post #15

Unsloth quants available: https://unsloth.ai/docs/models/qwen3.6

~/llama.cpp$ build-.../bin/llama-batched-bench -m models/....gguf -npp 512,1024,2048,4096,8192,16384,32768 -ntg 128 -npl 1 -c 36000

  On amd 7900xtx

  Qwen3.6-27B-Q4_K_M
  |    PP |     TG |    B |   N_KV |   T_PP s | S_PP t/s |   T_TG s | S_TG t/s |      T s |    S t/s |
  |-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
  |   512 |    128 |    1 |    640 |    0.743 |   689.35 |    4.605 |    27.80 |    5.348 |   119.68 |
  |  1024 |    128 |    1 |   1152 |    1.188 |   862.17 |    4.573 |    27.99 |    5.761 |   199.96 |
  |  2048 |    128 |    1 |   2176 |    2.566 |   798.09 |    4.602 |    27.81 |    7.168 |   303.57 |
  |  4096 |    128 |    1 |   4224 |    5.936 |   690.00 |    4.639 |    27.59 |   10.575 |   399.43 |
  |  8192 |    128 |    1 |   8320 |   15.034 |   544.90 |    4.729 |    27.06 |   19.763 |   420.98 |
  | 16384 |    128 |    1 |  16512 |   42.807 |   382.74 |    4.886 |    26.20 |   47.694 |   346.21 |
  | 32768 |    128 |    1 |  32896 |  137.377 |   238.53 |    5.188 |    24.67 |  142.566 |   230.74 |

  Qwen3.6-27B-IQ4_NL
  |    PP |     TG |    B |   N_KV |   T_PP s | S_PP t/s |   T_TG s | S_TG t/s |      T s |    S t/s |
  |-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
  |   512 |    128 |    1 |    640 |    0.535 |   957.45 |    3.715 |    34.45 |    4.250 |   150.59 |
  |  1024 |    128 |    1 |   1152 |    1.124 |   911.16 |    3.677 |    34.81 |    4.801 |   239.97 |
  |  2048 |    128 |    1 |   2176 |    2.447 |   836.89 |    3.698 |    34.62 |    6.145 |   354.13 |
  |  4096 |    128 |    1 |   4224 |    5.711 |   717.17 |    3.729 |    34.32 |    9.441 |   447.43 |
  |  8192 |    128 |    1 |   8320 |   14.615 |   560.52 |    3.821 |    33.50 |   18.436 |   451.30 |
  | 16384 |    128 |    1 |  16512 |   41.966 |   390.41 |    3.967 |    32.26 |   45.933 |   359.48 |
  | 32768 |    128 |    1 |  32896 |  135.789 |   241.32 |    4.253 |    30.09 |  140.042 |   234.90 |

  On mbp M2 Max

  Qwen3.6-27B-UD-Q8_K_XL
  |    PP |     TG |    B |   N_KV |   T_PP s | S_PP t/s |   T_TG s | S_TG t/s |      T s |    S t/s |
  |-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
  |   512 |    128 |    1 |    640 |    2.583 |   198.18 |   22.049 |     5.81 |   24.633 |    25.98 |
  |  1024 |    128 |    1 |   1152 |    8.321 |   123.06 |   22.364 |     5.72 |   30.685 |    37.54 |
  |  2048 |    128 |    1 |   2176 |   17.873 |   114.59 |   23.290 |     5.50 |   41.164 |    52.86 |
  |  4096 |    128 |    1 |   4224 |   41.967 |    97.60 |   23.624 |     5.42 |   65.591 |    64.40 |
  |  8192 |    128 |    1 |   8320 |   68.722 |   119.20 |   21.077 |     6.07 |   89.799 |    92.65 |
  | 16384 |    128 |    1 |  16512 |  142.184 |   115.23 |   22.026 |     5.81 |  164.210 |   100.55 |
  | 32768 |    128 |    1 |  32896 |  339.778 |    96.44 |   24.465 |     5.23 |  364.243 |    90.31 |

  Compared to similar prior models

  On amd 7900xtx

  Qwen3.6-35B-A3B-UD-Q4_K_S
  |    PP |     TG |    B |   N_KV |   T_PP s | S_PP t/s |   T_TG s | S_TG t/s |      T s |    S t/s |
  |-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
  |   512 |    128 |    1 |    640 |    0.203 |  2517.60 |    1.482 |    86.35 |    1.686 |   379.67 |
  |  1024 |    128 |    1 |   1152 |    0.427 |  2399.22 |    1.471 |    87.04 |    1.897 |   607.15 |
  |  2048 |    128 |    1 |   2176 |    0.946 |  2165.23 |    1.478 |    86.59 |    2.424 |   897.67 |
  |  4096 |    128 |    1 |   4224 |    2.253 |  1818.33 |    1.502 |    85.22 |    3.755 |  1125.01 |
  |  8192 |    128 |    1 |   8320 |    5.849 |  1400.51 |    1.525 |    83.91 |    7.375 |  1128.17 |
  | 16384 |    128 |    1 |  16512 |   17.115 |   957.27 |    1.589 |    80.55 |   18.705 |   882.78 |
  | 32768 |    128 |    1 |  32896 |   56.008 |   585.06 |    1.704 |    75.10 |   57.712 |   570.00 |

  Qwen3.6-35B-A3B-UD-IQ4_XS
  |    PP |     TG |    B |   N_KV |   T_PP s | S_PP t/s |   T_TG s | S_TG t/s |      T s |    S t/s |
  |-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
  |   512 |    128 |    1 |    640 |    0.204 |  2508.94 |    1.313 |    97.46 |    1.517 |   421.78 |
  |  1024 |    128 |    1 |   1152 |    0.423 |  2418.64 |    1.296 |    98.80 |    1.719 |   670.18 |
  |  2048 |    128 |    1 |   2176 |    0.946 |  2164.61 |    1.323 |    96.78 |    2.269 |   959.13 |
  |  4096 |    128 |    1 |   4224 |    2.235 |  1832.54 |    1.326 |    96.52 |    3.561 |  1186.06 |
  |  8192 |    128 |    1 |   8320 |    5.845 |  1401.44 |    1.352 |    94.70 |    7.197 |  1156.03 |
  | 16384 |    128 |    1 |  16512 |   17.096 |   958.38 |    1.417 |    90.33 |   18.513 |   891.94 |
  | 32768 |    128 |    1 |  32896 |   56.013 |   585.00 |    1.530 |    83.66 |   57.543 |   571.67 |

  Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_S
  |    PP |     TG |    B |   N_KV |   T_PP s | S_PP t/s |   T_TG s | S_TG t/s |      T s |    S t/s |
  |-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
  |   512 |    128 |    1 |    640 |    0.205 |  2499.78 |    1.483 |    86.31 |    1.688 |   379.16 |
  |  1024 |    128 |    1 |   1152 |    0.434 |  2361.36 |    1.448 |    88.40 |    1.882 |   612.25 |
  |  2048 |    128 |    1 |   2176 |    0.947 |  2161.87 |    1.478 |    86.62 |    2.425 |   897.27 |
  |  4096 |    128 |    1 |   4224 |    2.259 |  1813.00 |    1.472 |    86.94 |    3.732 |  1131.98 |
  |  8192 |    128 |    1 |   8320 |    5.892 |  1390.42 |    1.505 |    85.06 |    7.397 |  1124.85 |
  | 16384 |    128 |    1 |  16512 |   17.397 |   941.77 |    1.568 |    81.61 |   18.965 |   870.63 |
  | 32768 |    128 |    1 |  32896 |   56.296 |   582.07 |    1.690 |    75.74 |   57.986 |   567.31 |

  Nemotron-Cascade-2-30B-A3B-IQ4_XS
  |    PP |     TG |    B |   N_KV |   T_PP s | S_PP t/s |   T_TG s | S_TG t/s |      T s |    S t/s |
  |-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
  |   512 |    128 |    1 |    640 |    0.195 |  2622.33 |    0.972 |   131.69 |    1.167 |   548.30 |
  |  1024 |    128 |    1 |   1152 |    0.407 |  2514.76 |    0.934 |   137.10 |    1.341 |   859.16 |
  |  2048 |    128 |    1 |   2176 |    0.854 |  2396.99 |    0.942 |   135.90 |    1.796 |  1211.42 |
  |  4096 |    128 |    1 |   4224 |    1.895 |  2161.89 |    0.953 |   134.36 |    2.847 |  1483.50 |
  |  8192 |    128 |    1 |   8320 |    4.593 |  1783.70 |    0.967 |   132.43 |    5.559 |  1496.60 |
  | 16384 |    128 |    1 |  16512 |   12.213 |  1341.53 |    0.996 |   128.56 |   13.209 |  1250.10 |
  | 32768 |    128 |    1 |  32896 |   36.998 |   885.66 |    1.059 |   120.89 |   38.057 |   864.39 |

  On mbp M2 Max

  Qwen3.6-35B-A3B-UD-Q6_K_XL
  |    PP |     TG |    B |   N_KV |   T_PP s | S_PP t/s |   T_TG s | S_TG t/s |      T s |    S t/s |
  |-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
  |   512 |    128 |    1 |    640 |    0.540 |   947.31 |    2.489 |    51.42 |    3.030 |   211.22 |
  |  1024 |    128 |    1 |   1152 |    0.951 |  1077.21 |    3.237 |    39.54 |    4.188 |   275.10 |
  |  2048 |    128 |    1 |   2176 |    2.994 |   684.10 |    3.139 |    40.77 |    6.133 |   354.80 |
  |  4096 |    128 |    1 |   4224 |    6.245 |   655.85 |    3.210 |    39.88 |    9.455 |   446.75 |
  |  8192 |    128 |    1 |   8320 |   12.411 |   660.08 |    3.284 |    38.98 |   15.694 |   530.13 |
  | 16384 |    128 |    1 |  16512 |   28.321 |   578.51 |    3.584 |    35.71 |   31.905 |   517.53 |
  | 32768 |    128 |    1 |  32896 |   65.725 |   498.56 |    4.029 |    31.77 |   69.754 |   471.60 |

  Nemotron-Cascade-2-30B-A3B-Q8_0
  |    PP |     TG |    B |   N_KV |   T_PP s | S_PP t/s |   T_TG s | S_TG t/s |      T s |    S t/s |
  |-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
  |   512 |    128 |    1 |    640 |    0.528 |   969.13 |    2.036 |    62.87 |    2.564 |   249.59 |
  |  1024 |    128 |    1 |   1152 |    1.079 |   948.84 |    3.201 |    39.99 |    4.280 |   269.15 |
  |  2048 |    128 |    1 |   2176 |    3.390 |   604.10 |    2.952 |    43.36 |    6.342 |   343.11 |
  |  4096 |    128 |    1 |   4224 |    6.756 |   606.28 |    2.991 |    42.79 |    9.747 |   433.35 |
  |  8192 |    128 |    1 |   8320 |   13.647 |   600.30 |    3.061 |    41.81 |   16.708 |   497.97 |
  | 16384 |    128 |    1 |  16512 |   29.491 |   555.56 |    3.414 |    37.50 |   32.905 |   501.81 |
  | 32768 |    128 |    1 |  32896 |   65.867 |   497.49 |    3.663 |    34.95 |   69.530 |   473.12 |
Dang I saw some lowish numbers there for Spaks (and Strix). As I was eyeing a spark to get some CUDA exposure... :-O

Re: Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

#467
post #79

Earlier quoted context omitted.

What would be these additional vllm flags, if you don't mind sharing?

This is from an example from my Nomad cluster with two a5000's, which is a bit different what i have at work, but it will mostly apply to most modern 24G vram nvidia gpu. "--tensor-parallel-size", "2" - spread the LLM weights over 2 GPU's available "--max-model-len", "90000" - I've capped context window from ~256k to 90k. It allows us to have more concurrency and for our use cases it is enough. "--kv-cache-dtype", "f…

Thank you!

Re: Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

#468

Earlier quoted context omitted.

Getting ~36-33 tok/s (see the "S_TG t/s" column) on a 24GB Radeon RX 7900 XTX using llama.cpp's Vulkan backend: $ llama-server --version version: 8851 (e365e658f) $ llama-batched-bench -hf unsloth/Qwen3.6-27B-GGUF:IQ4_XS -npp 1000,2000,4000,8000,16000,32000 -ntg 128 -npl 1 -c 34000 | PP | TG | B | N_KV | T_PP s | S_PP t/s | T_TG s | S_TG t/s | T s | S t/s | |-------|--------|------|--------|----------|----------|----…

Did you try GPU/CPU mix with a bigger model?

Prompt processing is absolutely punishing:

    ./llama-batched-bench -hf unsloth/Qwen3.5-122B-A10B-GGUF:UD-IQ4_NL -npp 1000 -ntg 128 -npl 1 --cache-type-k q8_0 --cache-type-v q8_0 -c 18000 --n-cpu-moe 32
    |    PP |     TG |    B |   N_KV |   T_PP s | S_PP t/s |   T_TG s | S_TG t/s |      T s |    S t/s |
    |-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
    |  1000 |    128 |    1 |   1128 |   53.961 |    18.53 |    9.223 |    13.88 |   63.184 |    17.85 |

Re: Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

#470

Earlier quoted context omitted.

We made Unsloth Studio which should help :) 1. Auto best official parameters set for all models 2. Auto determines the largest quant that can fit on your PC / Mac etc 3. Auto determines max context length 4. Auto heals tool calls, provides python & bash + web search :)

Yea, I actually tried it out last time we had one of these threads. It's undeniably easy to use, but it is also very opinionated about things like the directory locations/layouts for various assets. I don't think I managed to get it to work with a simple flat directory full of pre-downloaded models on an NFS mount to my NAS. It also insists on re-downloading a 3GB model every time it is launches, even after I delete…

Oh my apologies I didn't respond - if only HN had a notifier haha

Oh yes we added a custom folder button which can pull .gguf files for now from any folder - it supports LM Studio and Ollama ones - but afreed it's still a mess.

One of the goals is to somehow quick search for .gguf folders, and add recommended folders - we currently have folders for Ollama and LM Studio for eg

Post reply on HN