base_model: orcarouter/Qwen3.8-27B-Uncensored
library_name: gguf
pipeline_tag: text-generation
tags:
- gguf
- qwen3.8
- quantization
- imatrix
- edit-preservation
language: [en, it, zh]
Qwen3.8-27B-Uncensored — Q38mix v6 (edit-preservation-aware GGUF)
A reasoned GGUF quantization of orcarouter/Qwen3.8-27B-Uncensored
(the rank-1 "OrcaRouter" abliterated Qwen3.8-27B, a hybrid gated-delta-net + attention architecture with a
native MTP head). 19.32 GB, 5.655 bpw. Fits an RTX 5090 (32 GB) with headroom for KV cache / long context.
What makes it different: to our knowledge it is the first GGUF quant of this model that is aware of the
abliteration edit itself and protects the tensors carrying it, while being tuned per-tensor against a true
BF16 reference (not a Q8 proxy).
Method (every choice measured)
- Custom importance matrix calibrated on this model: 861k tokens (agentic-coding + code + chat + Italian
- refusal-contrast), Qwen3 chat-template rendered, 125 chunks @ ctx 2048 with
--process-output.
- refusal-contrast), Qwen3 chat-template rendered, 125 chunks @ ctx 2048 with
- Per-tensor base allocation = Unsloth Dynamic V3 (extracted from
unsloth/Qwen3.8-27B-GGUFUD-Q4_K_XL) — full credit to
Unsloth for the base map. - Survival-guided edit protection: the 131 tensors carrying the shared rank-1 abliteration direction
(recovered first-hand, cosine 0.99999 across layers) are protected. - Measured reallocation: all
ffn_down→ Q6_K,attn_output→ Q6_K,token_embd→ Q6_K, driven by a
ΔKL-attribution sweep and validated against a BF16 reference.ssm_alpha/betakept Q8_0;conv1d/dt/aand
norms F32 (GDN control path, incompressible).lm_head/outputnot inflated (least sensitive — placebo).
Benchmarks — against a BF16 reference (the bias-free objective)
KL divergence vs the true Uncensored BF16 (held-out multi-domain corpus, 40×512 chunks). Reference is BF16, so
the comparison does not favour any build's Q8-kept tensors.
| quant | size | mean KL ↓ | same-top-1 ↑ | 99% KL tail ↓ | max KL ↓ |
|---|---|---|---|---|---|
| Q38mix v6 (this) | 19.32 GB | 0.008372 | 96.13 % | 0.0902 | 3.40 |
| OrcaRouter Q5_K_M (publisher's default) | 19.54 GB | 0.009791 | 96.03 % | 0.1082 | 4.21 |
| Unsloth-method + edit-protection (udprotected) | 19.09 GB | 0.008773 | 96.17 % | 0.0933 | 3.57 |
Honest reading:
- Beats the publisher's own OrcaRouter Q5_K_M on every measured axis — mean KL −14.5 %, tails −17 %/−19 %,
and it is smaller. This is a fair BF16-referenced comparison. - On par with the Unsloth method (udprotected): v6 edges mean KL and tails but is +0.23 GB larger, and
udprotected wins same-top-1 by 0.04 pp (noise). At equal size the two are a tie. Do not read this as
"beats Unsloth" — the differentiator over the Unsloth method is the edit protection, not the allocation. - Edit-survival (structural preservation of the abliteration direction): min ≈ 0.95, on par with OrcaRouter Q5
and far above an unprotected Unsloth-style quant (~0.86).
Trajectory fidelity (antirez ds4: greedy-LCP vs the Q8 reference's own greedy path, n=46)
| quant | greedy-LCP ↑ | exact-match ↑ |
|---|---|---|
| OrcaRouter Q5_K_M | 31.30 | 22 |
| Q38mix v6 (this) | 31.22 | 23 |
| Unsloth-method (udpure) | 30.61 | 22 |
v6 is co-leader on greedy-LCP (tied with OrcaRouter Q5 within noise) and best on exact-match, ahead of
the Unsloth method — i.e. it is strong on both the KL metric and the trajectory metric, which usually trade
off. (n=46; the ~0.6-LCP gaps are within the sign-test noise band — read as "competitive/co-leading", not
"decisively best".)
Honest caveats (read before relying on this)
- Reference is BF16 for the table above (the correct objective). An earlier Q8-referenced number for this
family was self-favourable and is not used here. - Behavioural refusal-rate (measured, full-window): on 25 AdvBench harmful + 15 Alpaca benign prompts,
n_predict=640 with the refusal judged on the answer after</think>(24/25 harmful outputs closed the
think block inside the window — a valid measurement, not the 128-token truncation the earlier bench suffered):
harmful-refusal 0.0, benign-refusal 0.067 — identical to the Q8 baseline. The abliteration is preserved
behaviourally; OrcaRouter Q5, by contrast, restored refusals (0.017 / 0.100). - Calibration/eval overlap: ~43 % of the KL-eval tokens share a source distribution with the calibration
corpus (documents disjoint, distribution not) — this favours any build using this imatrix over externally-built
quants. The numbers are honest but this is disclosed. - Built from the Q8_0 source with
--allow-requantize(a BF16-sourced rebuild is in progress and will supersede
this if measurably better).
Provenance
- GGUF sha256:
0e6f25cf9d203e0c75b9cb814e47346edc0d80b293d6a151287049da22ad4e04 - Engine: llama.cpp b10470 · quantized_from: Uncensored Q8_0 (
--allow-requantize) - Calibration: 861,596 tokens, 7 buckets · imatrix: 125 chunks @ ctx 2048,
--process-output - Full per-tensor map, imatrix, and calibration recipe: see
provenance.jsonand the accompanying files.
Variants in this repo
…Q38mix-v6.gguf(19.32 GB) — the flagship above (recommended).…Q38mix-v8-19.1G.gguf(19.14 GB, 5.60 bpw) — a slightly smaller variant (ffn_gate/up → IQ4_XS). BF16-ref
KL 0.008593, same-top-1 96.35 % (best of the set), tail (kld99) 0.098 — i.e. ≈v6 quality within noise,
−0.18 GB, with only the tail marginally worse. This is the mainline quality-per-byte frontier: a
size-matched reallocation (v6b) does not beat v6, and a real ~2 GB reduction is only reachable with the
ik_llama.cpp trellis fork (not servable in mainline llama.cpp / ollama / LM Studio). Full study:
the project'sdocs/quantization/andresearch/quant-obsessive/10-shrink-v6-executable.md.
Intended use
A high-fidelity local GGUF of an existing public abliterated model, for single-user coding-agent workloads on a
24–32 GB consumer GPU. It is a quantization of a model that is already publicly available; it adds no capability
beyond faithfully preserving that model at smaller size.