← back to catalog · registered 2026-08-22 13:56

hyperhuzaifa/Qwen3.6-27B-Uncensored-MTP-GGUF

hyperhuzaifa Qwen 27B GGUF multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/hyperhuzaifa%2FQwen3.6-27B-Uncensored-MTP-GGUF"
Response includes
  • classification m-uncensored
  • files 5
  • hub_downloads_all_time 8,604
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
9K
123 last 30d - cooling
Likes
4
Model age
3mo ago
created 2026-06-15
Downloads over time
Now8.6K→from362↑2,281%
03.1K6.3K9.4K362 on Jun 178.6K on Oct 118.6K on Oct 10JunJulAugSepOct
Jun 17 → Oct 11 · 56 snapshots · spans 116 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Quantizations
Q4_K
Tags
gguf qwen3.6 uncensored mtp speculative-decoding vision multimodal image-text-to-text en zh base_model:HauhauCS/Qwen3.6-27B-Uncensored-HauhauCS-Balanced base_model:merge:HauhauCS/Qwen3.6-27B-Uncensored-HauhauCS-Balanced

Related

Total size
16.6 GB
Files
5
Quantizations
3
Registered
2026-08-22 13:56
Last updated on HF
2026-06-15 18:14

Files by quantization

Q4_K 1 file 16.6 GB
Qwen3.6-27B-Uncensored-HauhauCS-Balanced-Q4_K_P-MTP.gguf 16.6 GB bdcbbf8a download
F16 1 file 885 MB
mmproj-Qwen3.6-27B-Uncensored-HauhauCS-Balanced-f16.gguf 885 MB 81c5dfa3 download
Auxiliary files 3 files 14.2 KB
chat_template_unsloth.jinja 8.02 KB 24209a1f download
README.md 4.47 KB 34ae5fb2 download
.gitattributes 1.67 KB 1b914411 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • HauhauCS/Qwen3.6-27B-Uncensored-HauhauCS-Balanced
  • unsloth/Qwen3.6-27B-MTP-GGUF
    base_model_relation: merge
    tags:
  • gguf
  • qwen3.6
  • uncensored
  • mtp
  • speculative-decoding
  • vision
  • multimodal
    language:
  • en
  • zh
    pipeline_tag: image-text-to-text

Qwen3.6-27B-Uncensored-HauhauCS-Balanced — Q4_K_P + MTP (GGUF)

This is HauhauCS/Qwen3.6-27B-Uncensored-HauhauCS-Balanced (Q4_K_P) with the Qwen 3.6 MTP "nextn" head grafted in, so it runs speculative decoding on llama.cpp out of the box — roughly 2× faster single-stream decode, with byte-identical output to the original (speculative decoding is lossless; the draft head only proposes tokens, the model verifies every one).

The upstream HauhauCS GGUFs ship without the MTP head, so there was no MTP-accelerated uncensored Qwen-27B available. This fills that gap.

What's different

upstream Q4_K_P this Q4_K_P-MTP
MTP / speculative decode ❌ none ✅ baked-in nextn head
Single-stream decode (4090, q8 KV, ctx 8K) 39.9 t/s 80.6 t/s (2.02×)
Draft acceptance — ~64%
Output quality identical identical (lossless)
Vision (mmproj) ✅ ✅

How it was made

The MTP head is an extra decoder block (blk.64.*, 15 tensors) plus two metadata keys (block_count → 65, nextn_predict_layers = 1). Those tensors were transplanted from unsloth/Qwen3.6-27B-MTP-GGUF (the only public source of the Qwen-27B MTP head) into the HauhauCS Q4_K_P file, with all other tensors/metadata copied faithfully. Because HauhauCS is a near-lossless abliteration of the same Qwen/Qwen3.6-27B base the head was trained on, draft acceptance stays high. Mixed per-tensor quant within a single GGUF is fully supported by llama.cpp.

Usage (llama.cpp)

llama-server -m Qwen3.6-27B-Uncensored-HauhauCS-Balanced-Q4_K_P-MTP.gguf \
  --mmproj mmproj-Qwen3.6-27B-Uncensored-HauhauCS-Balanced-f16.gguf \
  -ngl 99 --flash-attn on -c 65536 \
  --spec-type draft-mtp --spec-draft-n-max 3 \
  -ctk q8_0 -ctv q8_0 --jinja
  • MTP: --spec-type draft-mtp (the head is in the main GGUF — no --model-draft needed). Requires a build with the merged Qwen MTP path (mainline ≥ b9542).
  • Context: fits ~65K tokens at q8 KV on a single 24 GB GPU with the vision projector loaded.
  • Chat template: the embedded template is stock-official Qwen 3.6. For agentic/tool-use, the unsloth fixed template (chat_template_unsloth.jinja, included) is recommended — it removes two over-eager exceptions and fixes tool-call argument serialization. Pass it with --chat-template-file chat_template_unsloth.jinja.
  • Thinking model: pass enable_thinking: false (template kwarg) for short/structured outputs; inline /no_think is not honored.

ik_llama.cpp (alternative)

ik_llama.cpp runs the same baked-in MTP head, but with its own flag dialect — -fa 1 instead of --flash-attn on, and -mtp --draft-max 3 instead of --spec-type draft-mtp (no --model-draft; the head is in the GGUF):

./build/bin/llama-server \
  -m Qwen3.6-27B-Uncensored-HauhauCS-Balanced-Q4_K_P-MTP.gguf \
  --mmproj mmproj-Qwen3.6-27B-Uncensored-HauhauCS-Balanced-f16.gguf \
  -ngl 99 -fa 1 -c 65536 -ctk q8_0 -ctv q8_0 \
  -mtp --draft-max 3 \
  --jinja --host 127.0.0.1 --port 8080

This is a dense model, so MTP is a clear win on ik_llama too (~83 t/s measured on a 4090 for a comparable Qwen-27B-MTP build). Note: some ik_llama builds segfault on single-GPU + MTP — if so, run dual-GPU (-sm layer -ts 1,1) or fall back to mainline.

Files

  • Qwen3.6-27B-Uncensored-HauhauCS-Balanced-Q4_K_P-MTP.gguf — weights + grafted MTP head
  • mmproj-Qwen3.6-27B-Uncensored-HauhauCS-Balanced-f16.gguf — vision projector
  • chat_template_unsloth.jinja — recommended (bug-fixed) chat template

Credits & license

  • HauhauCS — the uncensored (abliterated) base weights.
  • unsloth — the Qwen 3.6 MTP head (donor) and the fixed chat template.
  • Qwen — the Qwen3.6-27B base model.

Apache-2.0, inherited from the upstream models. This is a redistribution with an added (lossless) MTP head — all model behavior and quality is HauhauCS's; only decode speed changes.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-15Add ik_llama.cpp usage section (-mtp --draft-max)171e1884.5 KB
    Loading...
  2. 2026-06-15Link both base models (HauhauCS + unsloth) via base_model_relation=merge for ...9885c4f3.6 KB
    Loading...
  3. 2026-06-15Set base_model_relation=quantized (GGUF + grafted MTP head, not a finetune)1707a263.6 KB
    Loading...
  4. 2026-06-15Add model cardfcc86993.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration