← back to catalog · registered 2026-09-28 21:57

DoktorMincs/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W4A16-AutoRound

DoktorMincs 27B multimodal second-order
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/DoktorMincs%2FQwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W4A16-AutoRound"
Response includes
  • classification m3
  • files 20
  • author_summary 6 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-09-28

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text autoround w4a16 int4 4-bit compressed-tensors vllm marlin mtp

Related

Total size
28.1 GB
Files
20
Quantizations
1
Registered
2026-09-28 21:57
Last updated on HF
2026-09-28 21:18

Files by quantization

Auxiliary files 20 files 28.1 GB
model-00001-of-00006.safetensors 4.99 GB b19ed17e download
model-00005-of-00006.safetensors 4.99 GB 03957ec6 download
model-00004-of-00006.safetensors 4.97 GB 68365ace download
model-00003-of-00006.safetensors 4.96 GB ff492d1a download
model-00002-of-00006.safetensors 4.94 GB db2b8d27 download
model-00006-of-00006.safetensors 3.23 GB a019f641 download
tokenizer.json 19.1 MB 87a7830d download
vocab.json 6.41 MB 0aa0ce06 download
model.safetensors.index.json 148 KB aecc3d50 download
tokenizer_config.json 16.0 KB 08f41ce2 download
LICENSE 11.1 KB d6456956 download
config.json 10.7 KB 46c67cfa download
chat_template.jinja 8.74 KB 5a39aa2b download
README.md 6.25 KB 9a8ed298 download
NOTICE 1.83 KB d9a2fa80 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 213 B d04042de download

README current version from Hugging Face


license: apache-2.0
base_model: DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU
base_model_relation: quantized
library_name: transformers
pipeline_tag: image-text-to-text
tags:

  • qwen3_5
  • autoround
  • w4a16
  • int4
  • 4-bit
  • compressed-tensors
  • vllm
  • marlin
  • mtp
  • vision
  • uncensored

Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU — W4A16 AutoRound (attention + GatedDeltaNet excluded)

4-bit weight-only quantization (W4A16, int4 symmetric, group_size 128) of
DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU,
produced with AutoRound (llm-compressor 0.14 + auto-round 0.15.1, 400 iterations, 256
calibration samples
), with the full-attention and GatedDeltaNet projections kept in BF16
because measurement showed they carry most of the quantization damage at 4 bits.

  • Quantized → compressed-tensors pack-quantized (uint4b8), served by vLLM's Marlin
    kernels: the MLP projections only (192 modules: mlp.{gate,up,down}_proj × 64 layers).
  • Kept in BF16 (bit-identical to the base model):
    • All full-attention projections — self_attn.{q,k,v,o}_proj, 64 modules
    • All GatedDeltaNet projections — linear_attn.{in_proj_qkv,in_proj_z,out_proj}, 144 modules
    • Vision tower (model.visual.*, 333 tensors) and MTP head (mtp.*, 15 tensors)
    • lm_head, embeddings, norms, conv1d, and the GatedDeltaNet in_proj_a/in_proj_b
  • Calibration: 256 samples × 2048 tokens from
    neuralmagic/LLM_compression_calibration
    with the model's own chat template.
  • Size: ~30.2 GB (base BF16 ~52 GB). 70% of the quantized parameters stay in 4 bits.

Quality

Perplexity on 50 held-out texts from wikitext-103, 16,007 evaluated tokens — deliberately not the
calibration set. Same texts, same tokenization, same method, both models served by vLLM. The BF16
baseline is 8.0880 (reproduced identically across every run of this series).

version perplexity delta
base BF16 8.0880 —
fully quantized W4A16 (nothing excluded) 8.7945 +8.74%
this build (attention + GatedDeltaNet in BF16) 8.3888 +3.72%

Where the 4-bit damage comes from

Each group of modules was restored to BF16 in turn — surgically, on top of the fully-quantized 4-bit
checkpoint, without re-quantizing — and the perplexity re-measured on the identical corpus:

group restored to BF16 modules perplexity delta share of the +8.74%
(none — fully quantized) 0 8.7945 +8.74% —
GatedDeltaNet projections 144 8.6289 +6.69% 2.05 pp
MLP 192 8.6115 +6.47% 2.27 pp
attention projections 64 8.5515 +5.73% 3.01 pp
attention + GatedDeltaNet 208 8.3888 +3.72% 5.02 pp

Attention is still the single largest contributor at 4 bits (3.01 pp), and no single group was
enough to reach the 5% band — attention alone leaves +5.73%. The attention + GatedDeltaNet pair is
what crosses it.

The MLP is not excluded on purpose: it holds 17.1 B of the 24.3 B quantized parameters (70%), so
restoring it to BF16 would produce a 42 GB checkpoint against 52 GB for plain BF16 — the
quantization would stop being a compression. Keeping it in 4 bits is what makes this build 30.2 GB.

Honest note on the numbering

This build is dominated by its 6-bit sibling in the same series
(…-W6A16-AutoRound, 27.6 GB, +2.40%): that one is smaller and more faithful, and being smaller
it also reads fewer bytes per token. So on size, fidelity and decode bandwidth together, the 6-bit
build wins.

What this checkpoint establishes is that +3.72% is reachable at 4 bits for this model — the
project's stated target of ≤5% is met — and it documents exactly which module groups the 4-bit error
lives in. If you are choosing one artifact to serve, choose the W6A16 build; if you specifically want
a 4-bit MLP for its kernel/throughput characteristics, this is the most faithful 4-bit option we
found.

Serving notes

The quantized weights are 4-bit compressed-tensors uint4b8, served by Marlin:

Using MarlinLinearKernel for CompressedTensorsWNA16

Attention and GatedDeltaNet projections run as ordinary BF16 GEMMs, which is what costs the bytes —
those 208 modules are 20.4 B parameters that stay at 2 bytes each.

Usage (vLLM ≥ 0.28)

vllm serve DoktorMincs/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W4A16-AutoRound \
  --tensor-parallel-size 2 \
  --max-model-len 32768 \
  --reasoning-parser qwen3 \
  --enable-prefix-caching \
  --speculative-config '{"method": "mtp", "num_speculative_tokens": 1}'
  • Vision inputs work normally (image/video); the vision tower runs in BF16.
  • The MTP draft head loads from the same checkpoint for speculative decoding.
  • Validated on vLLM 0.28: text, MTP speculative decoding and vision confirmed; 192 packed modules,
    208 weights left in BF16.
  • This is a thinking model: with a small max_tokens the whole budget can be consumed by the
    reasoning trace, leaving the visible answer empty. Use max_tokens ≥ 1024 and prefer the chat
    endpoint over raw completion.

Notes

  • Architecture: Qwen3_5ForConditionalGeneration (hybrid: 64 layers, 3:1 GatedDeltaNet
    linear-attention : full-attention, 27-block ViT, 1 MTP layer).
  • The exclusion is surgical: the MLP was AutoRound-tuned with the attention and GatedDeltaNet
    quantized, so a re-quantization with those excluded from the start would likely do slightly better.
  • No training, fine-tuning or abliteration was performed here — this is a weight-only quantization.
    The behavioural characteristics, abliteration and usage warnings of the upstream model are
    inherited unchanged.
  • Perplexity measures language-modelling fidelity only; it is not a substitute for task-specific
    evaluation.

License

Inherited from the upstream model, which declares the Apache License 2.0. See LICENSE and
NOTICE in this repository. Upstream chain: Qwen/Qwen3.8-27B → this checkpoint's base by DavidAU.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.