license: apache-2.0
base_model: Qwen/Qwen3.6-27B
tags:
- quantized
- gguf
- 3-bit
- qwen3.6
- aard-q3
- abliterated
- uncensored
model_type: qwen3_6
quantized_by: aaardpark
Qwen3.6-27B abliterated | aard-Q3
11 GB of Qwen 3.6-27B with the refusal direction surgically removed.
| Refusal rate (6 hard prompts) | KL vs base on harmless | |
|---|---|---|
| Base FP16 | 6 / 6 = 100% | 0 |
| abliterated FP16 | ~0 / 6 (1 reframe, 5 comply) | 0.0056 mean |
| aard-Q3 (this file) | inherits from FP16 | inherits |
For reference: the non-abliterated aard-Q3 hits 47/50 on GSM8K (94%). Abliteration is a rank-1 weight perturbation — capability hit is below quantization noise.
What "abliterated" means
Removed the refusal direction at layer 49 (cohen's d ≈ 7.0 between harmful and harmless prompt activations). One unit vector projected out of every weight that writes to the residual stream — embed_tokens, every block's o_proj / out_proj / down_proj (131 tensors total, ~5 GB of changes out of 54 GB).
Mean KL divergence vs base on 64 harmless Alpaca prompts: 0.0056 (Heretic's "clean" reference is ~0.08 — this is 14× under that).
Run it
huggingface-cli download aaardpark/Qwen3.6-27B-abliterated-GGUF \
qwen3.6-27B-abliterated-aaardpark-uniform-Q3_K.gguf --local-dir .
llama-cli -m qwen3.6-27B-abliterated-aaardpark-uniform-Q3_K.gguf -ngl 99 -c 32768
Same llama.cpp / runtime requirements as the non-abliterated version (build 8670+, thinking model, budget 2048+ tokens for hard reasoning).
Quick stats
| File | qwen3.6-27B-abliterated-aaardpark-uniform-Q3_K.gguf |
| Size | 11 GB |
| Format | GGUF, uniform Q3_K (3.59 BPW) |
| Refusal direction | layer 49, rank-1 ortho-projection |
| Min VRAM | 16 GB |
| Throughput | ~30 tok/s on Apple M-series |
| Native context | 262K |
More from aaardpark
- Qwen 3.6 27B (no abliteration) — 11 GB, 94% GSM8K
- Qwen 3.5 27B GGUF — 11 GB, 96% GSM8K
- gemma-4-31B-it — 15.3 GB, 96% GSM8K