license: apache-2.0
base_model: llmfan46/Qwen3.8-27B-Uncensored-Heretic-Native-MTP-Preserved
base_model_relation: quantized
pipeline_tag: text-generation
library_name: gguf
tags:
- gguf
- llama.cpp
- mtp
- speculative-decoding
- imatrix
- uncensored
- heretic
- qwen3.8
Qwen3.8-27B Uncensored Heretic — MTP, sized for 12 / 16 / 24 GB GPUs
GGUF quants of llmfan46/Qwen3.8-27B-Uncensored-Heretic-Native-MTP-Preserved,
built so the model and its native MTP head fit fully in VRAM on common consumer cards.
Each file applies Unsloth's per-tensor "UD" recipe for the base Qwen3.8-27B, copied tensor by tensor,
to llmfan46's weights. The MTP block (blk.64) stays at q6_K / q8_0 in every tier.
Every claim below was measured; the methodology is at the end.
Files
| File | Size | Target card | KLD vs Q8_0 (prose / code) |
|---|---|---|---|
Qwen3.8-27B-Uncensored-Heretic-MTP-UD-IQ3_XXS.gguf |
10.18 GiB | 12 GB | 0.090 / 0.059 |
Qwen3.8-27B-Uncensored-Heretic-MTP-UD-IQ4_XS.gguf |
13.27 GiB | 16 GB | 0.027 / 0.019 |
Qwen3.8-27B-Uncensored-Heretic-MTP-UD-Q5_K_XL.gguf |
19.44 GiB | 24 GB | 0.0045 / 0.0037 |
Quality against the other published quants of the same weights
Mean KL divergence against a Q8_0 reference of the same BF16 (lower is better), with the share of
tokens whose top-1 prediction matches the reference.
| Tier | Quant | GiB | KLD prose | KLD code | Same top-1 (prose / code) |
|---|---|---|---|---|---|
| 24 GB | this repo UD-Q5_K_XL | 19.44 | 0.0045 | 0.0037 | 97.2% / 98.5% |
| 16 GB | mradermacher i1-IQ4_XS ¹ | 14.26 | 0.0202 | 0.0149 | 94.3% / 96.8% |
| 16 GB | this repo UD-IQ4_XS | 13.27 | 0.0268 | 0.0192 | 93.2% / 96.5% |
| 16 GB | llmfan46 Q3_K_M | 13.48 | 0.0639 | 0.0456 | 89.3% / 95.0% |
| 16 GB | mradermacher i1-IQ3_M | 11.89 | 0.0649 | 0.0465 | 89.8% / 94.9% |
| 12 GB | this repo UD-IQ3_XXS | 10.18 | 0.0904 | 0.0587 | 87.4% / 94.5% |
| 12 GB | mradermacher i1-Q2_K | 10.12 | 0.1551 | 0.1017 | 82.9% / 92.6% |
¹ Better than UD-IQ4_XS, but at 14.26 GiB it does not leave room for MTP plus a useful context on a 16 GB card.
- 16 GB: 2.4× lower KLD than the two quants that fit with MTP (llmfan46 Q3_K_M, which is also 0.2 GiB larger, and mradermacher i1-IQ3_M).
- 12 GB: 1.7× lower KLD than mradermacher i1-Q2_K at the same size.
- 24 GB: effectively indistinguishable from Q8_0. No other 24 GB-class quant was measured, so no comparison is claimed.
Speed (RTX 4080 16 GB, measured)
llama.cpp b11457 (CUDA 13.4), MTP on, 2 draft tokens, ubatch 256, q4_0 KV cache, 32K context,
a 15,216-token prompt (a C++ source file), 600-token answers, mean of 3 fixed seeds.
| Quant | Code | Prose | MTP acceptance (code / prose) | Prompt eval |
|---|---|---|---|---|
| this repo UD-IQ4_XS | 55 tok/s (49–62) | 50 tok/s | 83% / 58% | ~1,360 tok/s |
| this repo UD-IQ3_XXS | 77 tok/s | 65 tok/s | 83% / 61% | ~1,440 tok/s |
| mradermacher i1-IQ3_M | 70 tok/s | 54 tok/s | 83% / 54% | ~1,480 tok/s |
| mradermacher i1-Q2_K | 72 tok/s | 59 tok/s | 79% / 53% | ~1,160 tok/s |
| llmfan46 Q3_K_M | 42 tok/s (32–50) | 41 tok/s | 84% / 61% | ~1,200 tok/s |
| UD-IQ4_XS without MTP (ubatch 512, unseeded) | 29 tok/s | 29 tok/s | — | ~1,500 tok/s |
- MTP roughly doubles decode speed at this depth. Two draft tokens beat three on the same seeds
(68 vs 56 tok/s code, 55 vs 36 prose, at 40K), because fewer drafts get rejected. - Keeping the MTP head at q6_K changes acceptance very little (83% vs 83% on code, 58% vs 54% on prose
against a quant with the head at IQ3_S). The quality gain of these files comes from the trunk, not the head. - The UD-IQ3_XXS numbers come from a 16 GB card with headroom. A 12 GB card has less memory bandwidth and will be slower.
Recommended settings
16 GB — UD-IQ4_XS
llama-server -m Qwen3.8-27B-Uncensored-Heretic-MTP-UD-IQ4_XS.gguf -ngl 99 --flash-attn on \
-c 32768 -ub 256 -ctk q4_0 -ctv q4_0 -np 1 \
--spec-type draft-mtp --spec-draft-n-max 2
This sits at the edge of 16 GB. Measured on the same card on different days, 40K of context gave
68, 51 and 42 tok/s on code with the desktop holding ~1.2, ~1.4 and ~1.9 GB of VRAM. When the model spills, Windows
moves part of it to system RAM ("shared GPU memory" in Task Manager) and decode drops. During one run, opening a
Chrome window mid-generation dropped decode to 0.2 tok/s. If you run a desktop on the same GPU, use 32K or less
and watch shared GPU memory: it should stay near its idle value. 48K with MTP produced the same tokens as 40K
and was 22% slower, from spill alone.
12 GB — UD-IQ3_XXS
llama-server -m Qwen3.8-27B-Uncensored-Heretic-MTP-UD-IQ3_XXS.gguf -ngl 99 --flash-attn on \
-c 16384 -ub 256 -ctk q4_0 -ctv q4_0 -np 1 \
--spec-type draft-mtp --spec-draft-n-max 2
Not run on a 12 GB card. The context is computed from the buffers llama.cpp reported at 32K: weights 10,020 MiB,
recurrent state 449 MiB, compute ~260 MiB, KV ~22 KB per token including the MTP draft cache, about 11.6 GB in
total at 32K. On a headless 12 GB card ~32K should fit. With a desktop on the same card, start at 16K.
24 GB — UD-Q5_K_XL
Same flags with -c 65536. Also computed rather than measured: ~22.4 GB at 64K.
Caveats
- Thinking. The chat template accepts
reasoning_effort=none,low/minimal,medium(default) orhigh/xhigh, pluschat_template_kwargs.enable_thinking. Thinking stays on when tools are present unlessauto_disable_thinking_with_toolsis set. - Uncensored. Refusal removal is llmfan46's work (Heretic v2, MPOA); their card reports 3/100 refusals and
KLD 0.0244 against the original Qwen3.8-27B. That was not re-measured here, and quantization can make behaviour
near the old refusal boundary less stable. The model will comply with requests the original refuses; you are
responsible for how you use it. - Vision. These files are text-only. llmfan46's
mmprojwas built for the same weights and should work
for image input, but it was not tested with these quants.
How these were made
- BF16 GGUF from llmfan46 (SHA-256
088e2801…5a062). - The per-tensor types of
unsloth/Qwen3.8-27B-GGUF(UD-IQ3_XXS,UD-IQ4_XS,UD-Q5_K_XL) read from the GGUF
headers. All 866 tensors match llmfan46's BF16 in name and shape. llama-quantize --tensor-type-filewith mradermacher's imatrix for these exact weights. Every recipe was
dry-run and checked tensor by tensor before quantizing. Final sizes match Unsloth's files.
One gotcha if you do this yourself: --tensor-type-file patterns are std::regex_search patterns. A plainoutput.weight=q5_k line also matches every blk.N.attn_output.weight and silently re-types them.
Anchor and escape each name: ^output\.weight$=q5_k.
Evaluation. Reference: Q8_0 from the same BF16. Text: wikitext-2 test and two llama.cpp source files
(~260 KB). llama-perplexity --kl-divergence, context 1024, 40 chunks per text (~20K scored tokens each,
standard error ±3%). Speed: llama-server as above, prompt from server-task.cpp. Scripts are in scripts/.
Credits
- Qwen for Qwen3.8-27B.
- llmfan46 for the Heretic weights with the MTP head preserved.
- Unsloth for the dynamic per-tensor recipes.
- mradermacher for the imatrix.
- llama.cpp.
License: Apache-2.0, as the base model.