← back to catalog · registered 2026-08-22 13:56

DoktorMincs/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-W4A16-AutoRound

DoktorMincs Qwen 24B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/DoktorMincs%2FQwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-W4A16-AutoRound"
Response includes
  • classification m3
  • files 15
  • hub_downloads_all_time 6,555
  • author_summary 6 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
7K
139 last 30d - cooling
Likes
2
Model age
2mo ago
created 2026-07-28
Downloads over time
Now6.6K→from0↑0%
02.4K4.9K7.3K0 on Jul 296.6K on Oct 11JulAugSepOct
Jul 29 → Oct 11 · 51 snapshots · spans 74 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
safetensors qwen3_5 autoround w4a16 gptq compressed-tensors vllm mtp vision base_model:DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP base_model:quantized:DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP license:other

Related

Total size
18.1 GB
Files
15
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-28 15:05

Files by quantization

Auxiliary files 15 files 18.1 GB
model-00003-of-00004.safetensors 5.00 GB bdd18c99 download
model-00002-of-00004.safetensors 4.99 GB fa9e8b41 download
model-00001-of-00004.safetensors 4.98 GB f66a1ae2 download
model-00004-of-00004.safetensors 3.15 GB 6175c20b download
tokenizer.json 19.1 MB 6f32ce20 download
model.safetensors.index.json 194 KB aa135eb0 download
config.json 15.7 KB 3aef5983 download
tokenizer_config.json 14.9 KB 5cc018ff download
chat_template-instruct.jinja 11.6 KB 53066082 download
chat_template.jinja 11.5 KB 82faea87 download
README.md 3.29 KB d5e73bb3 download
.gitattributes 1.53 KB 52373fe2 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 213 B d04042de download

README current version from Hugging Face


license: other
base_model: DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP
tags:

  • qwen3_5
  • autoround
  • w4a16
  • gptq
  • compressed-tensors
  • vllm
  • mtp
  • vision

Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP — W4A16 AutoRound

4-bit weight-only quantization (W4A16, int4 symmetric, group_size 128) of
DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP
using AutoRound via llm-compressor 0.12 + auto-round 0.13.

  • Quantized: all language-model Linear layers (400 modules: attention, MLP and GatedDeltaNet projections) → compressed-tensors pack-quantized (served by vLLM's Marlin kernels).
  • Preserved in BF16 (untouched, bit-identical to the base model):
    • Vision tower (model.visual.*, 333 tensors)
    • MTP / Multi-Token Prediction head (mtp.*, 15 tensors) — enables speculative decoding
    • lm_head, embeddings, norms, conv1d, and the GatedDeltaNet in_proj_a/in_proj_b projections (output dim 48 < group_size 128)
  • Calibration: 128 samples × 2048 tokens from neuralmagic/LLM_compression_calibration with the model's own chat template, 200 AutoRound iterations per layer.
  • Size: ~19.5 GB (base BF16: ~52 GB).

Usage (vLLM ≥ 0.26)

Production example (4× RTX 3090, tool calling + reasoning parser + MTP speculative decoding).
Replace the local path with the repo id DoktorMincs/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-W4A16-AutoRound to load directly from the Hub:

export CUDA_DEVICE_ORDER=PCI_BUS_ID
export OMP_NUM_THREADS=4
export CUDA_VISIBLE_DEVICES=0,1,2,3
export PYTORCH_ALLOC_CONF=expandable_segments:True
export NCCL_P2P_DISABLE=1
export FLASHINFER_DISABLE_VERSION_CHECK=1

exec vllm serve \
    /root/models/DoktorMincs/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-W4A16-AutoRound \
    --served-model-name "qwen3.6-27b" \
    --tensor-parallel-size 4 \
    --max-model-len 163840 \
    --override-generation-config '{"temperature": 0.6, "top_p": 0.95, "top_k": 20, "min_p": 0.0, "presence_penalty": 0.0, "repetition_penalty": 1.0}' \
    --tool-call-parser qwen3_coder \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --trust-remote-code \
    --enable-prefix-caching \
    --max-num-batched-tokens 8192 \
    --gpu-memory-utilization 0.88 \
    --disable-custom-all-reduce \
    --max-num-seqs 16 \
    --host 0.0.0.0 \
    --port 3434 \
    --speculative-config '{"method": "mtp", "num_speculative_tokens": 1}'
  • Vision inputs work normally (image/video); the vision tower runs in BF16.
  • The MTP draft head loads from the same checkpoint for speculative decoding.
  • Validated on vLLM 0.26 (text, MTP speculative decoding, vision).

Notes

  • Architecture: Qwen3_5ForConditionalGeneration (hybrid: 64 layers, 3:1 GatedDeltaNet linear-attention : full-attention, 27-block ViT, 1 MTP layer).
  • chat_template-instruct.jinja is the author's alternate template shipped with the base model.
  • Calibration data is general (math/code/logic/science QA); for heavy RP/creative use, a domain-matched calibration set may further improve fidelity.

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-28Add production vllm serve example (TP4, tool calling, MTP spec decode)e311f073.3 KB
    Loading...
  2. 2026-07-28Upload W4A16 AutoRound quantization (llm-compressor, MTP+vision BF16 preserved)9eb81502.3 KB
    Loading...
  3. 2026-07-28initial commitae1de9a28 B
    Loading...

Discussions 1 thread

  1. 2026-09-17Help Please!closed3 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration