license: apache-2.0
base_model: Qwen/Qwen3.6-35B-A3B
language:
- en
- zh
- multilingual
tags: - uncensored
- qwen3.6
- moe
- gguf
- mtp
- speculative-decoding
- vision
- multimodal
pipeline_tag: image-text-to-text
Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive — IQ4_XS + grafted MTP head
The HauhauCS uncensored finetune of Qwen3.6-35B-A3B, in IQ4_XS, with its
Multi-Token-Prediction (MTP) head fused into the same file so multi-token speculative
decoding works with no separate drafter model.
One file, nothing else to download, no second model competing for VRAM.
This is a derivative of Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive
by HauhauCS, built from theirIQ4_XSquant. All credit for the model and its
uncensoring goes to them — see Credits. This upload only adds the MTP head.
What's in the file
| Architecture | qwen35moe |
| Base | Qwen3.6-35B-A3B (Apache-2.0) |
| Finetune | Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive (Aggressive variant) |
| Quantization | IQ4_XS, imatrix (general.file_type = 30) |
| Layers | 40 trunk + 1 MTP block (block_count = 41, nextn_predict_layers = 1) |
| Experts | 256 per layer, 8 routed per token |
| Native context | 262,144 |
| Tensors | 753 (733 trunk + 20 MTP) |
| File size | 17.95 GiB (19,275,482,560 bytes / 19.3 GB) |
The 20 grafted blk.40.* tensors are ~521 MiB and kept at Q4_K_M from the MTP
sidecar. They are the only tensors not inherited from HauhauCS's IQ4_XS quant.
This file is the language model only. Vision needs HauhauCS'smmproj-Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-f16.gguf
alongside it, exactly as with the original IQ4_XS.
⚠️ Requirements
Your build must support the qwen35moe architecture and --spec-type draft-mtp.
Unsloth documents MTP running on
mainline llama.cpp for their own Qwen3.6 MTP
GGUFs, so a recent mainline build is the first thing to try — see their
MTP guide. This particular file was
built and measured on the buun-llama-cpp
fork, and its architecture string is qwen35moe; if your build does not recognise that
string, use one that does.
Without --spec-type draft-mtp the MTP block is skipped at load and the file behaves as an
ordinary 40-layer model.
Usage
llama-server \
-m Qwen3.6-35B-A3B-Uncensored-IQ4_XS-MTP.gguf \
-c 131072 \
-ctk turbo4 -ctv turbo4 -fa on \
-ngl 999 -ncmoe 36 -lm none -t 6 \
-b 3072 -ub 1536 -np 1 \
--moe-cache 1600 \
--spec-type draft-mtp --spec-draft-n-max 2 \
--temp 0.6 --top-p 0.95 --top-k 20 --min-p 0 \
--host 127.0.0.1 --port 8080 --jinja
No -md, no draft model path — the head is already in the file.
Add --mmproj mmproj-...-f16.gguf for vision — but note that vision and MTP cannot be
used together: --mmproj is not supported while MTP is enabled, and neither is-np > 1 (both per unsloth's notes). If you need vision, drop --spec-type draft-mtp.
Keep --jinja (as the original page notes). The sampling values above are the original's
coding / precise tasks preset, which is what was used for the benchmarks below; see
Recommended Settings for the other presets.
--spec-draft-n-max 2 is the balanced draft depth — the same default unsloth recommends.n=1 favours coding-style output, n=3 favours prose, and n≥4 loses on coding.
Measured performance
All figures from one machine — RTX 3060 Ti 8 GB, Ryzen 5 5600X, 32 GB DDR4-2133 —
ctx 131072, 256 greedy tokens, coding prompt 5,575 tokens / prose prompt 4,645 tokens.
The three-way comparison was measured in a single interleaved job, so its columns are
directly comparable.
Three-way comparison
| config | prefill t/s | decode coding | decode prose | peak VRAM |
|---|---|---|---|---|
| baseline llama.cpp (q8_0 KV) | 852.4 | 32.13 | 32.46 | 7617 MiB |
| buun-llama.cpp + turbo4 KV + MoE expert cache | 981.9 | 38.82 | 37.75 | 7477 MiB |
| + MTP | 937.8 | 45.73 | 44.81 | 7656 MiB |
| prefill | decode coding | decode prose | |
|---|---|---|---|
| turbo4 + expert cache vs baseline | +15.2% | +20.8% | +16.3% |
| + MTP vs baseline | +10.0% | +42.3% | +38.0% |
| MTP vs the same setup without MTP | −4.5% | +17.8% | +18.7% |
Note that roughly half the headline gain comes from the KV codec and expert cache, not
from the MTP head — the MTP head contributes the +17.8% / +18.7% row.
Draft acceptance with MTP enabled: 71.4% / 64.4% (coding), 84.7% / 60.2% (prose),
mean accepted run 2.20-2.68 tokens. The MTP head is the model's own trained next-token
head, which is why acceptance is high — and why it helps on ordinary open-ended generation
where a distilled drafter would not.
The −4.5% prefill is real: the verify batch costs a little prefill and buys a lot of decode.
Deep context (64K code prompt)
| depth | prefill | decode | peak VRAM | acceptance |
|---|---|---|---|---|
| 5,575 tok | 934.9 t/s | 45.6 t/s | 7620-7658 MiB | 65.9% |
| 65,712 tok | 809.9 t/s | 42.5 t/s | 7860 MiB | 77.5% |
Decode holds up at depth (−6.8%) and acceptance improves with context (65.9% → 77.5%),
because the MTP head predicts better with more context to condition on. Prefill costs 13.4%.
Peak VRAM at 64K is ~200 MiB above the shallow figure, because a long prefill saturates the
expert cache. Size your placement for the deep-context peak, not the idle number.
Draft depth
--spec-draft-n-max |
coding | prose |
|---|---|---|
| 1 | +17.8% | +10.8% |
| 2 | +12.9% | +29.8% |
| 3 | +5.0% | +33.4% |
| 4 | −2.1% | +19.6% |
| 5 | +3.9% | +18.5% |
Measured at ctx 8192 with all experts on CPU, against its own 27.44 / 27.11 t/s no-MTP
baseline. These percentages are not comparable to the ctx-131072 tables above, which
use a different placement (expert cache on, -ncmoe 36). The same n_max = 2 setting
measures +17.8% / +18.7% there rather than +12.9% / +29.8%, so the same setting does
not give the same percentage in every configuration. Use this table to rank the depths
against each other, and the tables above for the gain to actually expect.
How the MTP quant was chosen
Five quantizations of the MTP sidecar were benchmarked to pick the one grafted in here.
BF16 was excluded outright at 3.5 GB — an order of magnitude too large for a head that
sits in the same VRAM budget as the expert cache.
Measured at ctx 8192 with all experts on CPU (-ncmoe 40, expert cache off, n_max 2),
coding 5,575 tok and prose 4,645 tok. Baseline without MTP: 26.83 t/s coding, 28.50 t/s
prose.
| MTP quant | sidecar size | coding | Δ | prose | Δ | acceptance (coding / prose) | mean run | peak VRAM |
|---|---|---|---|---|---|---|---|---|
| Q8_0 | 1898 MB | 30.19 | +10.9% | 34.34 | +25.8% | 67.6% / 87.6% | 2.35 / 2.74 | 4337 MiB |
| Q6_K | 1469 MB | 30.50 | +12.0% | 34.36 | +25.9% | 67.6% / 87.6% | 2.35 / 2.74 | 4072 MiB |
| Q4_K_M | 1202 MB | 30.43 | +11.8% | 34.86 | +27.7% | 68.4% / 89.1% | 2.36 / 2.77 | 3885 MiB |
| IQ4_XS | 1096 MB | 30.56 | +12.3% | 33.94 | +24.3% | 68.4% / 86.1% | 2.36 / 2.71 | 3839 MiB |
| Q3_K_M | 1000 MB | 30.01 | +10.3% | 33.10 | +21.2% | 66.2% / 86.1% | 2.32 / 2.71 | 3747 MiB |
| BF16 | 3562 MB | not run | — | — | — | — | — | — |
The headline result: the MTP head's quality barely depends on its quantization. Every
quant from Q3_K_M to Q8_0 lands within ~2% on decode throughput and within ~2 percentage
points on acceptance. Quadrupling the head's size from 1000 MB to 3562 MB would buy
essentially nothing. That makes the decision almost purely a VRAM question, which is
convenient given the head shares an 8 GB budget with the KV cache and expert pool.
Why Q4_K_M:
- It measured best on prose (+27.7%, the strongest single result in the table) and
best on acceptance (89.1%), and was within noise of the best on coding. It is the
only quant that was top of either column while never being worse than ~0.4 t/s off the
best anywhere. - The head is small in absolute terms — 521 MiB after stripping the duplicated
embedding and LM head — so stepping up from IQ4_XS to Q4_K_M costs only ~46 MiB of VRAM
out of the 966-1012 MiB the head occupies. Trading 46 MiB for the best measured
acceptance is cheap. The same logic is what rules out Q8_0 and Q6_K: they cost +450 and
+187 MiB over Q4_K_M for no measurable gain. - Q3_K_M is where the head visibly degrades. It was the slowest on both prompts
(+10.3% / +21.2%) with the lowest acceptance (66.2%) — the first quant in the sweep to
cost something real. It saves only ~90 MiB versus IQ4_XS, so the floor is around 4 bits.
IQ4_XS was the runner-up and the choice came down to a near-tie: it was marginally best on
coding (+12.3%) and 46 MiB lighter, but weaker on prose acceptance (86.1% vs 89.1%). Q4_K_M
was taken for the prose result and the higher-fidelity head.
Two caveats on reading this table. The acceptance column is exact, not noisy — the
Q4_K_M figures reproduced bit-for-bit (147/215 and 163/183) when the sweep was re-run, so
differences of a few tokens are real. The throughput column is not — each cell is a
single 256-token generation, and this machine drifts up to ~5% between runs, so treat
differences under ~0.5 t/s as noise. In particular the IQ4_XS coding win is not
distinguishable from Q4_K_M.
VRAM guidance
On an 8 GB card the MTP block costs roughly 0.7-1.0 GB depending on placement, and MTP
consumes VRAM that would otherwise pay for expert-cache hits or GPU-resident layers. Two
settings mattered here:
- Use an explicit
--moe-cache <MiB>budget, noton. The pool is sized as
free VRAM − reserve, so on a busy desktoponsilently shrinks the pool instead of
failing — you lose cache hit rate and see no error. -ctkd/-ctvd turbo4quantizes the MTP block's own KV cache, saving ~380 MiB at
131k context versus f16, with no measurable change in acceptance.
At -ncmoe 36 with --moe-cache 1600 this peaked at 7843 MiB on a 64K prompt — about
349 MiB of headroom. If that is too tight, --moe-cache 1250 measured 7673 MiB
(519 MiB headroom) for ~2% less decode.
Specs
- 35B total parameters, ~3B active per forward pass (MoE)
- 256 experts, 8 routed per token
- Hybrid architecture: linear attention + full softmax attention (3:1 ratio)
- 40 layers + 1 MTP block
- 262K native context
- Natively multimodal (text, image, video) — requires the
mmprojfile - Based on Qwen/Qwen3.6-35B-A3B
Recommended Settings
From the official Qwen authors, as listed on the original model page:
Thinking mode (default):
- General:
temperature=1.0, top_p=0.95, top_k=20, min_p=0, presence_penalty=1.5 - Coding/precise tasks:
temperature=0.6, top_p=0.95, top_k=20, min_p=0, presence_penalty=0
Non-thinking mode:
- General:
temperature=0.7, top_p=0.8, top_k=20, min_p=0, presence_penalty=1.5 - Reasoning tasks:
temperature=1.0, top_p=1.0, top_k=40, min_p=0, presence_penalty=2.0
Important:
- Keep at least 128K context to preserve thinking capabilities
- Use
--jinjawith llama.cpp for proper chat template handling - Vision support requires the
mmprojfile alongside the main GGUF
How the graft was made
The MTP head is a single nextn transformer block. Two details matter if you reproduce
or modify this file:
block_countcounts the nextn blocks. In this codebasen_layer() = n_layer_all - n_layer_nextn, and MTP blocks load at indices[n_layer, n_layer_all). So a 40-layer trunk plus one MTP block must declareblock_count = 41with its tensors atblk.40.*. Declaringblock_count = 40makes
the loader look forblk.39.nextn.*and fail.- The MTP block reuses the trunk's embeddings and LM head. Only the 20
blk.40.*
tensors are added;token_embd.weight,output.weightandoutput_norm.weightare
not duplicated — the block falls back to the model's own, and duplicate tensor names
are invalid GGUF anyway.
19,264,492,032 bytes of tensor data plus a 19-byte alignment trailer accounts for the full
file size. Everything in the trunk — tensors, tokenizer, chat template and metadata — is
byte for byte HauhauCS's IQ4_XS quant.
Credits
- HauhauCS — the original uncensored finetune
Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive,
including the imatrix used for this quantization
(general.basename = KL0.0764,general.finetune = 3Ref). This upload is built on
theirIQ4_XSquant and would not exist without it. Their Discord:
https://discord.gg/SZ5vacTXYf - Qwen / Alibaba Cloud — the Qwen3.6-35B-A3B base model (Apache-2.0).
- spiritbuun — buun-llama-cpp. Its
qwen35moeMTP implementation, TurboQuant KV codecs (turbo4) and MoE expert cache are
what produce the numbers above; those are features of the runtime, not of this file. - unsloth — the MTP head weights come from
unsloth/Qwen3.6-35B-A3B-MTP-GGUF,
whoseQwen3.6-35B-A3B-MTPmodel this file'sblk.40.*block was cut from. See their
MTP guide and Discord
(https://discord.gg/unsloth). If you only want MTP, their GGUFs already ship it fused —
this upload exists for the HauhauCS uncensored finetune specifically.
Limitations
- These are one machine's numbers and depend heavily on having enough VRAM for the
expert cache and on CPU-side expert offload. Treat them as a shape, not a benchmark. - Greedy generation. Decode figures come from greedy sampling, where the MTP head
accepts the most. Real use attemp 0.6will accept somewhat fewer drafts, so these
numbers are optimistic. - Prefill is slower with MTP — roughly 4.5% — so for long-context ingestion with no
generation it is a net loss. - Acceptance is task-dependent. Measured here on code review and analytical prose.
Highly repetitive and highly novel text both behave differently. - The MTP block is Q4_K_M while the trunk is IQ4_XS, so this is not a uniform quant.
It is a small fraction of the weights. - Uncensored model. This is the Aggressive variant — see the original page for what
that entails, and use responsibly.