← back to catalog · registered 2026-08-22 13:56

kkuspa/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-NVFP4A16

kkuspa Qwen 19B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/kkuspa%2FQwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-NVFP4A16"
Response includes
  • classification m3
  • files 14
  • hub_downloads_all_time 12,223
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
12K
229 last 30d - cooling
Likes
9
Model age
2mo ago
created 2026-07-28
Downloads over time
Now12.3K→from0↑0%
04.5K9K13.5K0 on Jul 2912.3K on Oct 11JulAugSepOct
Jul 29 → Oct 11 · 51 snapshots · spans 74 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Quantizations
BF16
Tags
transformers safetensors qwen3_5 image-text-to-text qwen3.6 nvfp4 w4a16 compressed-tensors vllm quantized uncensored heretic

Related

Total size
26.6 GB
Files
14
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-07-31 19:04

Files by quantization

BF16 1 file 810 MB
model-mtp-bf16.safetensors 810 MB 3a02a15b download
Auxiliary files 13 files 25.8 GB
model.safetensors 25.8 GB f3eb528a download
tokenizer.json 19.1 MB 6f32ce20 download
model.safetensors.index.json 138 KB beded885 download
config.json 29.9 KB 8d8796f1 download
tokenizer_config.json 14.9 KB 5cc018ff download
chat_template-instruct.jinja 11.6 KB 53066082 download
chat_template.jinja 11.5 KB 82faea87 download
README.md 6.16 KB d7130698 download
.gitattributes 1.53 KB 52373fe2 download
recipe.yaml 552 B e47618ad download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 214 B 8a2e2eff download

README current version from Hugging Face


license: apache-2.0
base_model: DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: transformers
language:

  • en
  • zh
    tags:
  • qwen3.6
  • nvfp4
  • w4a16
  • compressed-tensors
  • vllm
  • quantized
  • uncensored
  • heretic
  • creative-writing
  • roleplaying
  • vision-language
  • mtp

Fable-Fusion 711 for vLLM (NVFP4A16)

This is the vLLM version of DavidAU's Qwen3.6-27B Fable-Fusion 711. The weights are FP4. Everything sensitive stays BF16. The repo is 28 GB instead of 56. It serves as an OpenAI-compatible endpoint on any NVIDIA GPU with about 32 GB of VRAM, Ampere or newer. Blackwell is not required.

Scheme NVFP4A16 weight-only, compressed-tensors format, native vLLM support
Source quantized directly from the BF16 safetensors. No GGUF step in the lineage.
Kept in BF16 all 48 linear-attention (DeltaNet) layers, the vision tower, the MTP drafter, lm_head
Size 27.7 GB weights (BF16 parent: 55.6 GB)
VRAM ~32 GB minimum; tested on 1x RTX PRO 6000 Blackwell (96 GB)
Context up to 262,144 tokens, VRAM permitting
Vision included, unquantized
MTP speculative decoding verified: 1.56x decode speedup at draft depth 5
FP8 KV cache calibrated k_scale/v_scale included; verified with --kv-cache-dtype fp8
License apache-2.0, inherited from the parent

Quickstart

vllm serve kkuspa/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-NVFP4A16 \
  --max-model-len 32768 --gpu-memory-utilization 0.85

Tested on vLLM 0.26.0. The server selects the Marlin NVFP4 kernel and warns that the GPU lacks native FP4 compute. The warning is expected: activations are BF16 by design, so the FP4 tensor-core path does not apply.

For faster decoding, enable the built-in MTP drafter:

vllm serve kkuspa/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-NVFP4A16   --max-model-len 32768 --gpu-memory-utilization 0.85   --speculative-config '{"method":"qwen3_5_mtp","num_speculative_tokens":5}'

DavidAU tuned this model for specific sampler settings. They translate to the OpenAI API as follows:

from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

# Thinking mode (default template): creative work, reasoning
r = client.chat.completions.create(
    model="kkuspa/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-NVFP4A16",
    messages=[{"role": "user", "content": "Open a noir story set in a lighthouse."}],
    temperature=1.0, top_p=0.95, max_tokens=2048,
    extra_body={"top_k": 20, "min_p": 0.0, "repetition_penalty": 1.0},
)

# Coding: temperature=0.6, same top_p and top_k
# Instruct mode: temperature=0.7, top_p=0.80, presence_penalty=1.5

The repo ships two chat templates. chat_template.jinja (thinking) is the default. For instruct-style responses without a thinking block, pass --chat-template pointing at chat_template-instruct.jinja at serve time.

This repo or the GGUF?

  • llama.cpp, LM Studio, Ollama, one user on a consumer card: use DavidAU's GGUF repo.
  • vLLM, concurrent requests, an OpenAI-compatible server, or 100k+ context: use this repo.

At the time of publication this is the only quant of Fable-Fusion 711 that stock vLLM can serve. The other safetensors quants of this model use a custom runtime format.

Fidelity

Measured with lm-evaluation-harness 0.4.12 (vLLM backend, 0-shot, batch auto) against the BF16 parent on the same hardware:

Task BF16 NVFP4A16 Delta
ARC-Challenge (acc_norm) 0.6126 0.6203 +0.008
HellaSwag (acc_norm) 0.8502 0.8490 -0.001
Winogrande (acc) 0.7806 0.7814 +0.001

Deltas sit inside the standard error on every task (ARC-Challenge stderr is 0.014). The quantization is a data-free RTN pass. No calibration dataset was used, so nothing was tuned toward any benchmark. Absolute scores depend on harness settings; the point of this table is the like-for-like comparison, both configs measured 0-shot on the same harness and hardware.

Speed

Single request, 512-token greedy generations, 1x RTX PRO 6000 Blackwell (96 GB), vLLM 0.26.0:

Config Decode speed
Standard decoding 56.2 tok/s
MTP speculative, depth 5 87.4 tok/s (1.56x)

Acceptance across the test: 996 of 3000 drafted tokens (about 1.7 extra tokens per draft window). Throughput under concurrent load will differ.

What was quantized

llm-compressor 0.12.0, one-shot:

QuantizationModifier(
    targets="Linear",
    scheme="NVFP4A16",
    ignore=["lm_head", "re:.*visual.*", "re:.*linear_attn.*", "re:.*mtp.*"],
)

Only the Linear weights in the full-attention and MLP blocks carry FP4. The DeltaNet linear-attention layers, the vision tower, the MTP drafter, and lm_head keep BF16, because hybrid-attention layers and drafters degrade badly under 4-bit quantization. The exact recipe ships in recipe.yaml.

Packaging notes

  • The parent repo ships no image-processor configs. preprocessor_config.json and video_preprocessor_config.json here come from Qwen/Qwen3.6-27B.
  • The 15 MTP drafter tensors live in model-mtp-bf16.safetensors and are wired into model.safetensors.index.json. Transformers drops them on load; they were re-extracted from the parent checkpoint so speculative decoding remains possible.

Content notice

The parent model is refusal-ablated ("Uncensored", "Heretic") and tuned for unconstrained creative writing. This quant changes none of that. It will produce content that aligned models refuse. You are responsible for how you deploy it.

Caveats

  • The vision path loads but has not been benchmarked.

Lineage and attribution

Qwen/Qwen3.6-27B by the Qwen team, then the Fable-Fusion 711 multi-stage tune and merge by DavidAU with collaborators credited on the parent card, then this quantization. Thanks to all of them. Apache-2.0 throughout.

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-31FP8 KV cache calibration: static k/v scales (512-sample creative-writing cali...11f9ee06.2 KB
    Loading...
  2. 2026-07-28Upload README.md with huggingface_hube5e15b36.1 KB
    Loading...
  3. 2026-07-28Upload folder using huggingface_hub2d240845.5 KB
    Loading...

Discussions 1 thread

  1. 2026-07-31FP8 KV cache calibrationclosed10 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration