← back to catalog · registered 2026-09-30 22:58

Digitals06/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Agressive-MTP-GGUF

Digitals06 Qwen 35B GGUF MoE multimodal
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Digitals06%2FQwen3.6-35B-A3B-Uncensored-HauhauCS-Agressive-MTP-GGUF"
Response includes
  • classification m-uncensored
  • files 2
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
1
Model age
today
created 2026-09-30
Downloads over time
Now0→from0↑0%
00110 on Sep 300 on Oct 1SepOct
Sep 30 → Oct 1 · 2 snapshots · spans 1 day

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh multilingual
Tags
gguf uncensored qwen3.6 moe mtp speculative-decoding vision multimodal image-text-to-text en zh multilingual

Related

Total size
0 B
Files
2
Quantizations
1
Registered
2026-09-30 22:58
Last updated on HF
2026-09-30 23:30

Files by quantization

Auxiliary files 2 files 12.5 KB
README.md 11.0 KB 0fceb814 download
.gitattributes 1.48 KB a6344aac download

README current version from Hugging Face


license: apache-2.0
base_model: Qwen/Qwen3.6-35B-A3B
language:

  • en
  • zh
  • multilingual
    tags:
  • uncensored
  • qwen3.6
  • moe
  • gguf
  • mtp
  • speculative-decoding
  • vision
  • multimodal
    pipeline_tag: image-text-to-text

Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive — IQ4_XS + grafted MTP head

The HauhauCS uncensored finetune of Qwen3.6-35B-A3B, in IQ4_XS, with its
Multi-Token-Prediction (MTP) head fused into the same file so multi-token speculative
decoding works with no separate drafter model.

One file, nothing else to download, no second model competing for VRAM.

This is a derivative of Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive
by HauhauCS, built from their IQ4_XS quant. All credit for the model and its
uncensoring goes to them — see Credits. This upload only adds the MTP head.

What's in the file

Architecture qwen35moe
Base Qwen3.6-35B-A3B (Apache-2.0)
Finetune Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive (Aggressive variant)
Quantization IQ4_XS, imatrix (general.file_type = 30)
Layers 40 trunk + 1 MTP block (block_count = 41, nextn_predict_layers = 1)
Experts 256 per layer, 8 routed per token
Native context 262,144
Tensors 753 (733 trunk + 20 MTP)
File size 17.95 GiB (19,275,482,560 bytes / 19.3 GB)

The 20 grafted blk.40.* tensors are ~521 MiB and kept at Q4_K_M from the MTP
sidecar. They are the only tensors not inherited from HauhauCS's IQ4_XS quant.

This file is the language model only. Vision needs HauhauCS's
mmproj-Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-f16.gguf
alongside it, exactly as with the original IQ4_XS.

⚠️ Requirements

Your build must support the qwen35moe architecture and --spec-type draft-mtp.
Unsloth documents MTP running on
mainline llama.cpp for their own Qwen3.6 MTP
GGUFs, so a recent mainline build is the first thing to try — see their
MTP guide. This particular file was
built and measured on the buun-llama-cpp
fork, and its architecture string is qwen35moe; if your build does not recognise that
string, use one that does.

Without --spec-type draft-mtp the MTP block is skipped at load and the file behaves as an
ordinary 40-layer model.

Usage

llama-server \
  -m Qwen3.6-35B-A3B-Uncensored-IQ4_XS-MTP.gguf \
  -c 131072 \
  -ctk turbo4 -ctv turbo4 -fa on \
  -ngl 999 -ncmoe 36 -lm none -t 6 \
  -b 3072 -ub 1536 -np 1 \
  --moe-cache 1600 \
  --spec-type draft-mtp --spec-draft-n-max 2 \
  --temp 0.6 --top-p 0.95 --top-k 20 --min-p 0 \
  --host 127.0.0.1 --port 8080 --jinja

No -md, no draft model path — the head is already in the file.

Add --mmproj mmproj-...-f16.gguf for vision — but note that vision and MTP cannot be
used together
: --mmproj is not supported while MTP is enabled, and neither is
-np > 1 (both per unsloth's notes). If you need vision, drop --spec-type draft-mtp.
Keep --jinja (as the original page notes). The sampling values above are the original's
coding / precise tasks preset, which is what was used for the benchmarks below; see
Recommended Settings for the other presets.

--spec-draft-n-max 2 is the balanced draft depth — the same default unsloth recommends.
n=1 favours coding-style output, n=3 favours prose, and n≥4 loses on coding.

Measured performance

All figures from one machine — RTX 3060 Ti 8 GB, Ryzen 5 5600X, 32 GB DDR4-2133 —
ctx 131072, 256 greedy tokens, coding prompt 5,575 tokens / prose prompt 4,645 tokens.
The three-way comparison was measured in a single interleaved job, so its columns are
directly comparable.

Three-way comparison

config prefill t/s decode coding decode prose peak VRAM
baseline llama.cpp (q8_0 KV) 852.4 32.13 32.46 7617 MiB
llama.cpp + turbo4 KV + MoE expert cache 981.9 38.82 37.75 7477 MiB
this model + MTP 937.8 45.73 44.81 7656 MiB
prefill decode coding decode prose
turbo4 + expert cache vs baseline +15.2% +20.8% +16.3%
this model + MTP vs baseline +10.0% +42.3% +38.0%
MTP vs the same setup without MTP −4.5% +17.8% +18.7%

Note that roughly half the headline gain comes from the KV codec and expert cache, not
from the MTP head — the MTP head contributes the +17.8% / +18.7% row.

Draft acceptance with MTP enabled: 71.4% / 64.4% (coding), 84.7% / 60.2% (prose),
mean accepted run 2.20-2.68 tokens. The MTP head is the model's own trained next-token
head, which is why acceptance is high — and why it helps on ordinary open-ended generation
where a distilled drafter would not.

The −4.5% prefill is real: the verify batch costs a little prefill and buys a lot of decode.

Deep context (64K code prompt)

depth prefill decode peak VRAM acceptance
5,575 tok 934.9 t/s 45.6 t/s 7620-7658 MiB 65.9%
65,712 tok 809.9 t/s 42.5 t/s 7860 MiB 77.5%

Decode holds up at depth (−6.8%) and acceptance improves with context (65.9% → 77.5%),
because the MTP head predicts better with more context to condition on. Prefill costs 13.4%.

Peak VRAM at 64K is ~200 MiB above the shallow figure, because a long prefill saturates the
expert cache. Size your placement for the deep-context peak, not the idle number.

Draft depth

--spec-draft-n-max coding prose
1 +17.8% +10.8%
2 +12.9% +29.8%
3 +5.0% +33.4%
4 −2.1% +19.6%
5 +3.9% +18.5%

VRAM guidance

On an 8 GB card the MTP block costs roughly 0.7-1.0 GB depending on placement, and MTP
consumes VRAM that would otherwise pay for expert-cache hits or GPU-resident layers. Two
settings mattered here:

  • Use an explicit --moe-cache <MiB> budget, not on. The pool is sized as
    free VRAM − reserve, so on a busy desktop on silently shrinks the pool instead of
    failing — you lose cache hit rate and see no error.
  • -ctkd/-ctvd turbo4 quantizes the MTP block's own KV cache, saving ~380 MiB at
    131k context versus f16, with no measurable change in acceptance.

At -ncmoe 36 with --moe-cache 1600 this peaked at 7843 MiB on a 64K prompt — about
349 MiB of headroom. If that is too tight, --moe-cache 1250 measured 7673 MiB
(519 MiB headroom) for ~2% less decode.

Specs

  • 35B total parameters, ~3B active per forward pass (MoE)
  • 256 experts, 8 routed per token
  • Hybrid architecture: linear attention + full softmax attention (3:1 ratio)
  • 40 layers + 1 MTP block
  • 262K native context
  • Natively multimodal (text, image, video) — requires the mmproj file
  • Based on Qwen/Qwen3.6-35B-A3B

Recommended Settings

From the official Qwen authors, as listed on the original model page:

Thinking mode (default):

  • General: temperature=1.0, top_p=0.95, top_k=20, min_p=0, presence_penalty=1.5
  • Coding/precise tasks: temperature=0.6, top_p=0.95, top_k=20, min_p=0, presence_penalty=0

Non-thinking mode:

  • General: temperature=0.7, top_p=0.8, top_k=20, min_p=0, presence_penalty=1.5
  • Reasoning tasks: temperature=1.0, top_p=1.0, top_k=40, min_p=0, presence_penalty=2.0

Important:

  • Keep at least 128K context to preserve thinking capabilities
  • Use --jinja with llama.cpp for proper chat template handling
  • Vision support requires the mmproj file alongside the main GGUF

How the graft was made

The MTP head is a single nextn transformer block. Two details matter if you reproduce
or modify this file:

  1. block_count counts the nextn blocks. In this codebase
    n_layer() = n_layer_all - n_layer_nextn, and MTP blocks load at indices
    [n_layer, n_layer_all). So a 40-layer trunk plus one MTP block must declare
    block_count = 41 with its tensors at blk.40.*. Declaring block_count = 40 makes
    the loader look for blk.39.nextn.* and fail.
  2. The MTP block reuses the trunk's embeddings and LM head. Only the 20 blk.40.*
    tensors are added; token_embd.weight, output.weight and output_norm.weight are
    not duplicated — the block falls back to the model's own, and duplicate tensor names
    are invalid GGUF anyway.

19,264,492,032 bytes of tensor data plus a 19-byte alignment trailer accounts for the full
file size. Everything in the trunk — tensors, tokenizer, chat template and metadata — is
byte for byte HauhauCS's IQ4_XS quant.

Credits

  • HauhauCS — the original uncensored finetune
    Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive,
    including the imatrix used for this quantization
    (general.basename = KL0.0764, general.finetune = 3Ref). This upload is built on
    their IQ4_XS quant and would not exist without it. Their Discord:
    https://discord.gg/SZ5vacTXYf
  • Qwen / Alibaba Cloud — the Qwen3.6-35B-A3B base model (Apache-2.0).
  • spiritbuun — buun-llama-cpp. Its
    qwen35moe MTP implementation, TurboQuant KV codecs (turbo4) and MoE expert cache are
    what produce the numbers above; those are features of the runtime, not of this file.
  • unsloth — the MTP head weights come from
    unsloth/Qwen3.6-35B-A3B-MTP-GGUF,
    whose Qwen3.6-35B-A3B-MTP model this file's blk.40.* block was cut from. See their
    MTP guide and Discord
    (https://discord.gg/unsloth). If you only want MTP, their GGUFs already ship it fused —
    this upload exists for the HauhauCS uncensored finetune specifically.

Limitations

  • These are one machine's numbers and depend heavily on having enough VRAM for the
    expert cache and on CPU-side expert offload. Treat them as a shape, not a benchmark.
  • Greedy generation. Decode figures come from greedy sampling, where the MTP head
    accepts the most. Real use at temp 0.6 will accept somewhat fewer drafts, so these
    numbers are optimistic.
  • Prefill is slower with MTP — roughly 4.5% — so for long-context ingestion with no
    generation it is a net loss.
  • Acceptance is task-dependent. Measured here on code review and analytical prose.
    Highly repetitive and highly novel text both behave differently.
  • The MTP block is Q4_K_M while the trunk is IQ4_XS, so this is not a uniform quant.
    It is a small fraction of the weights.
  • Uncensored model. This is the Aggressive variant — see the original page for what
    that entails, and use responsibly.
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.