← back to catalog · registered 2026-08-22 13:56

zeeksa/Qwen3.6-27B-Uncensored-HauhauCS-Aggressive-DFlash-GGUF

zeeksa Qwen 27B GGUF second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/zeeksa%2FQwen3.6-27B-Uncensored-HauhauCS-Aggressive-DFlash-GGUF"
Response includes
  • classification m-uncensored
  • files 18
  • hub_downloads_all_time 9,599
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
10K
626 last 30d - cooling
Likes
4
Model age
2mo ago
created 2026-07-16
Downloads over time
Now9.8K→from1.3K↑633%
9104.1K7.4K10.6K1.3K on Jul 159.8K on Oct 11JulAugSepOct
Jul 15 → Oct 11 · 53 snapshots · spans 88 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Quantizations
BF16 IQ2 IQ3 IQ4 Q2_K Q3_K Q4_K Q5_K Q6_K Q8_0 Q8_K
Tags
gguf llama.cpp speculative-decoding dflash qwen3.6 uncensored not-for-all-audiences base_model:HauhauCS/Qwen3.6-27B-Uncensored-HauhauCS-Aggressive base_model:quantized:HauhauCS/Qwen3.6-27B-Uncensored-HauhauCS-Aggressive license:mit endpoints_compatible region:us

Related

Total size
166 GB
Files
18
Quantizations
13
Registered
2026-08-22 13:56
Last updated on HF
2026-07-16 15:31

Files by quantization

Q8_K 1 file 29.8 GB
Qwen3.6-27B-Uncensored-HauhauCS-Aggressive-Q8_K_P.gguf 29.8 GB 77672ff2 download
Q6_K 2 files 22.9 GB
Qwen3.6-27B-Uncensored-HauhauCS-Aggressive-Q6_K_P.gguf 21.6 GB f17b88ee download
Qwen3.6-27B-DFlash-Q6_K.gguf 1.33 GB 016d6969 download
Q5_K 2 files 20.5 GB
Qwen3.6-27B-Uncensored-HauhauCS-Aggressive-Q5_K_P.gguf 19.4 GB 9b123d76 download
Qwen3.6-27B-DFlash-Q5_K.gguf 1.14 GB db2bbbb2 download
Q4_K 2 files 17.3 GB
Qwen3.6-27B-Uncensored-HauhauCS-Aggressive-Q4_K_P.gguf 16.3 GB e44827d0 download
Qwen3.6-27B-DFlash-Q4_K_M.gguf 985 MB 71362369 download
IQ4 1 file 14.0 GB
Qwen3.6-27B-Uncensored-HauhauCS-Aggressive-IQ4_XS.gguf 14.0 GB 8d9c7934 download
Q3_K 1 file 13.3 GB
Qwen3.6-27B-Uncensored-HauhauCS-Aggressive-Q3_K_P.gguf 13.3 GB d94c86ae download
IQ3 2 files 22.9 GB
Qwen3.6-27B-Uncensored-HauhauCS-Aggressive-IQ3_M.gguf 11.7 GB 91164182 download
Qwen3.6-27B-Uncensored-HauhauCS-Aggressive-IQ3_XS.gguf 11.1 GB 96aac50c download
Q2_K 1 file 10.7 GB
Qwen3.6-27B-Uncensored-HauhauCS-Aggressive-Q2_K_P.gguf 10.7 GB f35137af download
IQ2 1 file 9.32 GB
Qwen3.6-27B-Uncensored-HauhauCS-Aggressive-IQ2_M.gguf 9.32 GB 284ceed3 download
BF16 1 file 3.23 GB
Qwen3.6-27B-DFlash-BF16.gguf 3.23 GB 39dd44b7 download
Q8_0 1 file 1.72 GB
Qwen3.6-27B-DFlash-Q8_0.gguf 1.72 GB 23b6c8eb download
F16 1 file 885 MB
mmproj-Qwen3.6-27B-Uncensored-HauhauCS-Aggressive-f16.gguf 885 MB 082ca68e download
Auxiliary files 2 files 6.49 KB
README.md 3.71 KB dc8e984b download
.gitattributes 2.78 KB 95fcb1e7 download

README current version from Hugging Face


license: mit
base_model: HauhauCS/Qwen3.6-27B-Uncensored-HauhauCS-Aggressive
tags:

  • gguf
  • llama.cpp
  • speculative-decoding
  • dflash
  • qwen3.6
  • uncensored
  • not-for-all-audiences

Qwen3.6-27B Uncensored HauhauCS Aggressive + DFlash — complete GGUF bundle

Everything needed to run
HauhauCS/Qwen3.6-27B-Uncensored-HauhauCS-Aggressive
with DFlash speculative decoding in llama.cpp: all target quants, the mmproj (vision),
and the DFlash draft head in five quant levels.

DFlash is lossless — output is byte-identical to running the target alone; the draft only
makes generation faster (measured +9–17% over the MTP head, see below). The target and
draft are always two separate files loaded together; any target quant pairs with any
draft quant.

Target model quants (choose one)

File Size
...-IQ2_M.gguf 10.0 GB
...-Q2_K_P.gguf 11.5 GB
...-IQ3_XS.gguf 12.0 GB
...-IQ3_M.gguf 12.6 GB
...-Q3_K_P.gguf 14.3 GB
...-IQ4_XS.gguf 15.1 GB
...-Q4_K_P.gguf 17.5 GB
...-Q5_K_P.gguf 20.8 GB
...-Q6_K_P.gguf 23.2 GB
...-Q8_K_P.gguf 32.0 GB

Plus mmproj-...-f16.gguf (0.93 GB) for vision input.

DFlash draft head quants (choose one)

File Size Notes
Qwen3.6-27B-DFlash-BF16.gguf 3.5 GB Best acceptance — recommended
Qwen3.6-27B-DFlash-Q8_0.gguf 1.9 GB Near-lossless
Qwen3.6-27B-DFlash-Q6_K.gguf 1.4 GB
Qwen3.6-27B-DFlash-Q5_K.gguf 1.2 GB
Qwen3.6-27B-DFlash-Q4_K_M.gguf 1.0 GB Smallest

Pairing: any head works with any target

The draft head and target quant are independent choices — llama.cpp loads them as two
separate models, so every head in this repo pairs with every target quant above. There is
no need for a Q2/Q3/IQ head to match a Q2/Q3/IQ target. Suggested pairing: BF16 head
if you have ~3.5 GB VRAM to spare (best acceptance), Q4_K_M head (1 GB) if squeezed.
Avoid quantizing the head harder than that: draft quality sets your speedup (we measured
code acceptance drop 69% -> 56% just going BF16 -> Q8_0), so a Q2-class head would erase
the benefit of speculation.

Usage (llama.cpp with DFlash support, arch dflash)

llama-server \
  -m Qwen3.6-27B-Uncensored-HauhauCS-Aggressive-Q8_K_P.gguf \
  --mmproj mmproj-Qwen3.6-27B-Uncensored-HauhauCS-Aggressive-f16.gguf \
  -ngl 99 -fa on \
  --spec-type draft-dflash \
  --spec-draft-model Qwen3.6-27B-DFlash-BF16.gguf \
  --spec-draft-n-max 6 \
  --spec-draft-ngl 99

Measured results

Q8_K_P target + BF16 draft, RTX PRO 6000 Blackwell 96 GB, 200-token greedy completions.
Baseline: same target with the Qwen3.6-27B MTP head (--spec-type draft-mtp, n-max 5).

Workload MTP baseline DFlash Speedup DFlash acceptance (mean run len)
Code 108.5 tok/s 126.6 tok/s +16.6% 69.3% (5.10)
Math 146.0 tok/s 159.9 tok/s +9.5% 85.6% (6.03)
Chat 88.4 tok/s 96.7 tok/s +9.5% 43.9% (3.60)

Q8_0 draft on code: 98 tok/s at 55.6% acceptance — quantizing the draft trades acceptance
for VRAM; prefer BF16 unless memory-constrained. The drafter was trained against stock
Qwen3.6-27B yet acceptance holds up well on this finetune; output verified identical with
and without speculation.

Attribution

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-16Upload README.md with huggingface_hub4bb11df3.7 KB
    Loading...
  2. 2026-07-16Upload README.md with huggingface_hub9c3adaa3.1 KB
    Loading...
  3. 2026-07-16Upload folder using huggingface_hub52d68b82.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration