← back to catalog · registered 2026-10-09 21:58

codavidgarcia/Qwen3.8-27B-Uncensored-Heretic-MTP-UD-GGUF

codavidgarcia 27B GGUF second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/codavidgarcia%2FQwen3.8-27B-Uncensored-Heretic-MTP-UD-GGUF"
Response includes
  • classification m3
  • files 5
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-09

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Quantizations
IQ3 IQ4 Q5_K
Tags
gguf llama.cpp mtp speculative-decoding imatrix uncensored heretic qwen3.8 text-generation base_model:llmfan46/Qwen3.8-27B-Uncensored-Heretic-Native-MTP-Preserved base_model:quantized:llmfan46/Qwen3.8-27B-Uncensored-Heretic-Native-MTP-Preserved license:apache-2.0
Total size
42.9 GB
Files
5
Quantizations
4
Registered
2026-10-09 21:58
Last updated on HF
2026-10-09 21:32

Files by quantization

Q5_K 1 file 19.4 GB
Qwen3.8-27B-Uncensored-Heretic-MTP-UD-Q5_K_XL.gguf 19.4 GB 29c1ac57 download
IQ4 1 file 13.3 GB
Qwen3.8-27B-Uncensored-Heretic-MTP-UD-IQ4_XS.gguf 13.3 GB 7fe255b9 download
IQ3 1 file 10.2 GB
Qwen3.8-27B-Uncensored-Heretic-MTP-UD-IQ3_XXS.gguf 10.2 GB 8c9dcd8e download
Auxiliary files 2 files 9.23 KB
README.md 7.49 KB cbd42584 download
.gitattributes 1.74 KB 95e981f4 download

README current version from Hugging Face


license: apache-2.0
base_model: llmfan46/Qwen3.8-27B-Uncensored-Heretic-Native-MTP-Preserved
base_model_relation: quantized
pipeline_tag: text-generation
library_name: gguf
tags:

  • gguf
  • llama.cpp
  • mtp
  • speculative-decoding
  • imatrix
  • uncensored
  • heretic
  • qwen3.8

Qwen3.8-27B Uncensored Heretic — MTP, sized for 12 / 16 / 24 GB GPUs

GGUF quants of llmfan46/Qwen3.8-27B-Uncensored-Heretic-Native-MTP-Preserved,
built so the model and its native MTP head fit fully in VRAM on common consumer cards.
Each file applies Unsloth's per-tensor "UD" recipe for the base Qwen3.8-27B, copied tensor by tensor,
to llmfan46's weights. The MTP block (blk.64) stays at q6_K / q8_0 in every tier.

Every claim below was measured; the methodology is at the end.

Files

File Size Target card KLD vs Q8_0 (prose / code)
Qwen3.8-27B-Uncensored-Heretic-MTP-UD-IQ3_XXS.gguf 10.18 GiB 12 GB 0.090 / 0.059
Qwen3.8-27B-Uncensored-Heretic-MTP-UD-IQ4_XS.gguf 13.27 GiB 16 GB 0.027 / 0.019
Qwen3.8-27B-Uncensored-Heretic-MTP-UD-Q5_K_XL.gguf 19.44 GiB 24 GB 0.0045 / 0.0037

Quality against the other published quants of the same weights

Mean KL divergence against a Q8_0 reference of the same BF16 (lower is better), with the share of
tokens whose top-1 prediction matches the reference.

Tier Quant GiB KLD prose KLD code Same top-1 (prose / code)
24 GB this repo UD-Q5_K_XL 19.44 0.0045 0.0037 97.2% / 98.5%
16 GB mradermacher i1-IQ4_XS ¹ 14.26 0.0202 0.0149 94.3% / 96.8%
16 GB this repo UD-IQ4_XS 13.27 0.0268 0.0192 93.2% / 96.5%
16 GB llmfan46 Q3_K_M 13.48 0.0639 0.0456 89.3% / 95.0%
16 GB mradermacher i1-IQ3_M 11.89 0.0649 0.0465 89.8% / 94.9%
12 GB this repo UD-IQ3_XXS 10.18 0.0904 0.0587 87.4% / 94.5%
12 GB mradermacher i1-Q2_K 10.12 0.1551 0.1017 82.9% / 92.6%

¹ Better than UD-IQ4_XS, but at 14.26 GiB it does not leave room for MTP plus a useful context on a 16 GB card.

  • 16 GB: 2.4× lower KLD than the two quants that fit with MTP (llmfan46 Q3_K_M, which is also 0.2 GiB larger, and mradermacher i1-IQ3_M).
  • 12 GB: 1.7× lower KLD than mradermacher i1-Q2_K at the same size.
  • 24 GB: effectively indistinguishable from Q8_0. No other 24 GB-class quant was measured, so no comparison is claimed.

Speed (RTX 4080 16 GB, measured)

llama.cpp b11457 (CUDA 13.4), MTP on, 2 draft tokens, ubatch 256, q4_0 KV cache, 32K context,
a 15,216-token prompt (a C++ source file), 600-token answers, mean of 3 fixed seeds.

Quant Code Prose MTP acceptance (code / prose) Prompt eval
this repo UD-IQ4_XS 55 tok/s (49–62) 50 tok/s 83% / 58% ~1,360 tok/s
this repo UD-IQ3_XXS 77 tok/s 65 tok/s 83% / 61% ~1,440 tok/s
mradermacher i1-IQ3_M 70 tok/s 54 tok/s 83% / 54% ~1,480 tok/s
mradermacher i1-Q2_K 72 tok/s 59 tok/s 79% / 53% ~1,160 tok/s
llmfan46 Q3_K_M 42 tok/s (32–50) 41 tok/s 84% / 61% ~1,200 tok/s
UD-IQ4_XS without MTP (ubatch 512, unseeded) 29 tok/s 29 tok/s — ~1,500 tok/s
  • MTP roughly doubles decode speed at this depth. Two draft tokens beat three on the same seeds
    (68 vs 56 tok/s code, 55 vs 36 prose, at 40K), because fewer drafts get rejected.
  • Keeping the MTP head at q6_K changes acceptance very little (83% vs 83% on code, 58% vs 54% on prose
    against a quant with the head at IQ3_S). The quality gain of these files comes from the trunk, not the head.
  • The UD-IQ3_XXS numbers come from a 16 GB card with headroom. A 12 GB card has less memory bandwidth and will be slower.

Recommended settings

16 GB — UD-IQ4_XS

llama-server -m Qwen3.8-27B-Uncensored-Heretic-MTP-UD-IQ4_XS.gguf -ngl 99 --flash-attn on \
  -c 32768 -ub 256 -ctk q4_0 -ctv q4_0 -np 1 \
  --spec-type draft-mtp --spec-draft-n-max 2

This sits at the edge of 16 GB. Measured on the same card on different days, 40K of context gave
68, 51 and 42 tok/s on code with the desktop holding ~1.2, ~1.4 and ~1.9 GB of VRAM. When the model spills, Windows
moves part of it to system RAM ("shared GPU memory" in Task Manager) and decode drops. During one run, opening a
Chrome window mid-generation dropped decode to 0.2 tok/s. If you run a desktop on the same GPU, use 32K or less
and watch shared GPU memory: it should stay near its idle value. 48K with MTP produced the same tokens as 40K
and was 22% slower, from spill alone.

12 GB — UD-IQ3_XXS

llama-server -m Qwen3.8-27B-Uncensored-Heretic-MTP-UD-IQ3_XXS.gguf -ngl 99 --flash-attn on \
  -c 16384 -ub 256 -ctk q4_0 -ctv q4_0 -np 1 \
  --spec-type draft-mtp --spec-draft-n-max 2

Not run on a 12 GB card. The context is computed from the buffers llama.cpp reported at 32K: weights 10,020 MiB,
recurrent state 449 MiB, compute ~260 MiB, KV ~22 KB per token including the MTP draft cache, about 11.6 GB in
total at 32K. On a headless 12 GB card ~32K should fit. With a desktop on the same card, start at 16K.

24 GB — UD-Q5_K_XL

Same flags with -c 65536. Also computed rather than measured: ~22.4 GB at 64K.

Caveats

  • Thinking. The chat template accepts reasoning_effort = none, low/minimal, medium (default) or
    high/xhigh, plus chat_template_kwargs.enable_thinking. Thinking stays on when tools are present unless
    auto_disable_thinking_with_tools is set.
  • Uncensored. Refusal removal is llmfan46's work (Heretic v2, MPOA); their card reports 3/100 refusals and
    KLD 0.0244 against the original Qwen3.8-27B. That was not re-measured here, and quantization can make behaviour
    near the old refusal boundary less stable. The model will comply with requests the original refuses; you are
    responsible for how you use it.
  • Vision. These files are text-only. llmfan46's mmproj was built for the same weights and should work
    for image input, but it was not tested with these quants.

How these were made

  1. BF16 GGUF from llmfan46 (SHA-256 088e2801…5a062).
  2. The per-tensor types of unsloth/Qwen3.8-27B-GGUF (UD-IQ3_XXS, UD-IQ4_XS, UD-Q5_K_XL) read from the GGUF
    headers. All 866 tensors match llmfan46's BF16 in name and shape.
  3. llama-quantize --tensor-type-file with mradermacher's imatrix for these exact weights. Every recipe was
    dry-run and checked tensor by tensor before quantizing. Final sizes match Unsloth's files.

One gotcha if you do this yourself: --tensor-type-file patterns are std::regex_search patterns. A plain
output.weight=q5_k line also matches every blk.N.attn_output.weight and silently re-types them.
Anchor and escape each name: ^output\.weight$=q5_k.

Evaluation. Reference: Q8_0 from the same BF16. Text: wikitext-2 test and two llama.cpp source files
(~260 KB). llama-perplexity --kl-divergence, context 1024, 40 chunks per text (~20K scored tokens each,
standard error ±3%). Speed: llama-server as above, prompt from server-task.cpp. Scripts are in scripts/.

Credits

License: Apache-2.0, as the base model.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration