Earlier quoted context omitted.
This is the best I got, all with Unsloth's quantizations. Laguna-S-2.1:UD-Q4_K_XL (no MTP) pp=186.4 t/s tg=27.8 t/s Qwen3.6-35B:UD-Q4_K_XL (with MTP) pp=404.4 t/s tg=83.2 t/s Qwen3.6-27B:UD-Q4_K_XL (recorded pre-MTP) pp=343 t/s tg=12.1 t/s Laguna actually performed better than I remembered. I thought it was slower.
Have you benchmarked against full precision models for accuracy/ performance?
Here's a summary of what I have for 27B. I used unsloth's UD-Q{3-6}_K_XL quants across 11 evals. The values are pretty linear between Q3 and Q6.
Qwen3.6-27B Q3 Q6
ARC-Challenge 97.0 97.0
BIG-Bench Hard 57.9 59.3
GPQA Diamond 77.8 83.3
GSM8K 92.4 92.6
Hendrycks Math 35.5 38.9
HumanEval 80.5 85.4
HumanEval+ 75.0 79.3
IFEval 87.3 88.0
MBPP 75.2 77.2
MBPP+ 88.4 88.9
MMLU-Pro 83.1 83.5