license: apache-2.0
base_model: Qwen/Qwen3.6-35B-A3B
language:
- en
- zh
- multilingual
tags: - uncensored
- qwen3.6
- moe
- gguf
- mtp
- speculative-decoding
- vision
- multimodal
pipeline_tag: image-text-to-text
Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive — IQ4_XS + grafted MTP head
The HauhauCS uncensored finetune of Qwen3.6-35B-A3B, in IQ4_XS, with its
Multi-Token-Prediction (MTP) head fused into the same file so multi-token speculative
decoding works with no separate drafter model.
One file, nothing else to download, no second model competing for VRAM.
This is a derivative of Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive
by HauhauCS, built from theirIQ4_XSquant. All credit for the model and its
uncensoring goes to them — see Credits. This upload only adds the MTP head.
What's in the file
| Architecture | qwen35moe |
| Base | Qwen3.6-35B-A3B (Apache-2.0) |
| Finetune | Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive (Aggressive variant) |
| Quantization | IQ4_XS, imatrix (general.file_type = 30) |
| Layers | 40 trunk + 1 MTP block (block_count = 41, nextn_predict_layers = 1) |
| Experts | 256 per layer, 8 routed per token |
| Native context | 262,144 |
| Tensors | 753 (733 trunk + 20 MTP) |
| File size | 17.95 GiB (19,275,482,560 bytes / 19.3 GB) |
The 20 grafted blk.40.* tensors are ~521 MiB and kept at Q4_K_M from the MTP
sidecar. They are the only tensors not inherited from HauhauCS's IQ4_XS quant.
This file is the language model only. Vision needs HauhauCS'smmproj-Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-f16.gguf
alongside it, exactly as with the original IQ4_XS.
⚠️ Requirements
Your build must support the qwen35moe architecture and --spec-type draft-mtp.
Unsloth documents MTP running on
mainline llama.cpp for their own Qwen3.6 MTP
GGUFs, so a recent mainline build is the first thing to try — see their
MTP guide. This particular file was
built and measured on the buun-llama-cpp
fork, and its architecture string is qwen35moe; if your build does not recognise that
string, use one that does.
Without --spec-type draft-mtp the MTP block is skipped at load and the file behaves as an
ordinary 40-layer model.
Usage
llama-server \
-m Qwen3.6-35B-A3B-Uncensored-IQ4_XS-MTP.gguf \
-c 131072 \
-ctk turbo4 -ctv turbo4 -fa on \
-ngl 999 -ncmoe 36 -lm none -t 6 \
-b 3072 -ub 1536 -np 1 \
--moe-cache 1600 \
--spec-type draft-mtp --spec-draft-n-max 2 \
--temp 0.6 --top-p 0.95 --top-k 20 --min-p 0 \
--host 127.0.0.1 --port 8080 --jinja
No -md, no draft model path — the head is already in the file.
Add --mmproj mmproj-...-f16.gguf for vision — but note that vision and MTP cannot be
used together: --mmproj is not supported while MTP is enabled, and neither is-np > 1 (both per unsloth's notes). If you need vision, drop --spec-type draft-mtp.
Keep --jinja (as the original page notes). The sampling values above are the original's
coding / precise tasks preset, which is what was used for the benchmarks below; see
Recommended Settings for the other presets.
--spec-draft-n-max 2 is the balanced draft depth — the same default unsloth recommends.n=1 favours coding-style output, n=3 favours prose, and n≥4 loses on coding.
Measured performance
All figures from one machine — RTX 3060 Ti 8 GB, Ryzen 5 5600X, 32 GB DDR4-2133 —
ctx 131072, 256 greedy tokens, coding prompt 5,575 tokens / prose prompt 4,645 tokens.
The three-way comparison was measured in a single interleaved job, so its columns are
directly comparable.
Three-way comparison
| config | prefill t/s | decode coding | decode prose | peak VRAM |
|---|---|---|---|---|
| baseline llama.cpp (q8_0 KV) | 852.4 | 32.13 | 32.46 | 7617 MiB |
| llama.cpp + turbo4 KV + MoE expert cache | 981.9 | 38.82 | 37.75 | 7477 MiB |
| this model + MTP | 937.8 | 45.73 | 44.81 | 7656 MiB |
| prefill | decode coding | decode prose | |
|---|---|---|---|
| turbo4 + expert cache vs baseline | +15.2% | +20.8% | +16.3% |
| this model + MTP vs baseline | +10.0% | +42.3% | +38.0% |
| MTP vs the same setup without MTP | −4.5% | +17.8% | +18.7% |
Note that roughly half the headline gain comes from the KV codec and expert cache, not
from the MTP head — the MTP head contributes the +17.8% / +18.7% row.
Draft acceptance with MTP enabled: 71.4% / 64.4% (coding), 84.7% / 60.2% (prose),
mean accepted run 2.20-2.68 tokens. The MTP head is the model's own trained next-token
head, which is why acceptance is high — and why it helps on ordinary open-ended generation
where a distilled drafter would not.
The −4.5% prefill is real: the verify batch costs a little prefill and buys a lot of decode.
Deep context (64K code prompt)
| depth | prefill | decode | peak VRAM | acceptance |
|---|---|---|---|---|
| 5,575 tok | 934.9 t/s | 45.6 t/s | 7620-7658 MiB | 65.9% |
| 65,712 tok | 809.9 t/s | 42.5 t/s | 7860 MiB | 77.5% |
Decode holds up at depth (−6.8%) and acceptance improves with context (65.9% → 77.5%),
because the MTP head predicts better with more context to condition on. Prefill costs 13.4%.
Peak VRAM at 64K is ~200 MiB above the shallow figure, because a long prefill saturates the
expert cache. Size your placement for the deep-context peak, not the idle number.
Draft depth
--spec-draft-n-max |
coding | prose |
|---|---|---|
| 1 | +17.8% | +10.8% |
| 2 | +12.9% | +29.8% |
| 3 | +5.0% | +33.4% |
| 4 | −2.1% | +19.6% |
| 5 | +3.9% | +18.5% |
VRAM guidance
On an 8 GB card the MTP block costs roughly 0.7-1.0 GB depending on placement, and MTP
consumes VRAM that would otherwise pay for expert-cache hits or GPU-resident layers. Two
settings mattered here:
- Use an explicit
--moe-cache <MiB>budget, noton. The pool is sized as
free VRAM − reserve, so on a busy desktoponsilently shrinks the pool instead of
failing — you lose cache hit rate and see no error. -ctkd/-ctvd turbo4quantizes the MTP block's own KV cache, saving ~380 MiB at
131k context versus f16, with no measurable change in acceptance.
At -ncmoe 36 with --moe-cache 1600 this peaked at 7843 MiB on a 64K prompt — about
349 MiB of headroom. If that is too tight, --moe-cache 1250 measured 7673 MiB
(519 MiB headroom) for ~2% less decode.
Specs
- 35B total parameters, ~3B active per forward pass (MoE)
- 256 experts, 8 routed per token
- Hybrid architecture: linear attention + full softmax attention (3:1 ratio)
- 40 layers + 1 MTP block
- 262K native context
- Natively multimodal (text, image, video) — requires the
mmprojfile - Based on Qwen/Qwen3.6-35B-A3B
Recommended Settings
From the official Qwen authors, as listed on the original model page:
Thinking mode (default):
- General:
temperature=1.0, top_p=0.95, top_k=20, min_p=0, presence_penalty=1.5 - Coding/precise tasks:
temperature=0.6, top_p=0.95, top_k=20, min_p=0, presence_penalty=0
Non-thinking mode:
- General:
temperature=0.7, top_p=0.8, top_k=20, min_p=0, presence_penalty=1.5 - Reasoning tasks:
temperature=1.0, top_p=1.0, top_k=40, min_p=0, presence_penalty=2.0
Important:
- Keep at least 128K context to preserve thinking capabilities
- Use
--jinjawith llama.cpp for proper chat template handling - Vision support requires the
mmprojfile alongside the main GGUF
How the graft was made
The MTP head is a single nextn transformer block. Two details matter if you reproduce
or modify this file:
block_countcounts the nextn blocks. In this codebasen_layer() = n_layer_all - n_layer_nextn, and MTP blocks load at indices[n_layer, n_layer_all). So a 40-layer trunk plus one MTP block must declareblock_count = 41with its tensors atblk.40.*. Declaringblock_count = 40makes
the loader look forblk.39.nextn.*and fail.- The MTP block reuses the trunk's embeddings and LM head. Only the 20
blk.40.*
tensors are added;token_embd.weight,output.weightandoutput_norm.weightare
not duplicated — the block falls back to the model's own, and duplicate tensor names
are invalid GGUF anyway.
19,264,492,032 bytes of tensor data plus a 19-byte alignment trailer accounts for the full
file size. Everything in the trunk — tensors, tokenizer, chat template and metadata — is
byte for byte HauhauCS's IQ4_XS quant.
Credits
- HauhauCS — the original uncensored finetune
Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive,
including the imatrix used for this quantization
(general.basename = KL0.0764,general.finetune = 3Ref). This upload is built on
theirIQ4_XSquant and would not exist without it. Their Discord:
https://discord.gg/SZ5vacTXYf - Qwen / Alibaba Cloud — the Qwen3.6-35B-A3B base model (Apache-2.0).
- spiritbuun — buun-llama-cpp. Its
qwen35moeMTP implementation, TurboQuant KV codecs (turbo4) and MoE expert cache are
what produce the numbers above; those are features of the runtime, not of this file. - unsloth — the MTP head weights come from
unsloth/Qwen3.6-35B-A3B-MTP-GGUF,
whoseQwen3.6-35B-A3B-MTPmodel this file'sblk.40.*block was cut from. See their
MTP guide and Discord
(https://discord.gg/unsloth). If you only want MTP, their GGUFs already ship it fused —
this upload exists for the HauhauCS uncensored finetune specifically.
Limitations
- These are one machine's numbers and depend heavily on having enough VRAM for the
expert cache and on CPU-side expert offload. Treat them as a shape, not a benchmark. - Greedy generation. Decode figures come from greedy sampling, where the MTP head
accepts the most. Real use attemp 0.6will accept somewhat fewer drafts, so these
numbers are optimistic. - Prefill is slower with MTP — roughly 4.5% — so for long-context ingestion with no
generation it is a net loss. - Acceptance is task-dependent. Measured here on code review and analytical prose.
Highly repetitive and highly novel text both behave differently. - The MTP block is Q4_K_M while the trunk is IQ4_XS, so this is not a uniform quant.
It is a small fraction of the weights. - Uncensored model. This is the Aggressive variant — see the original page for what
that entails, and use responsibly.