← back to catalog · registered 2026-08-22 13:56

tacodevs/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-FP8

tacodevs Qwen 24B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/tacodevs%2FQwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-FP8"
Response includes
  • classification m3
  • files 10
  • hub_downloads_all_time 369
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
369
177 last 30d - stable
Likes
0
Model age
2mo ago
created 2026-07-31
Downloads over time
Now449→from2↑22,350%
01653294942 on Jul 29449 on Oct 11JulAugSepOct
Jul 29 → Oct 11 · 51 snapshots · spans 74 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text fp8 compressed-tensors quantized vision-language multimodal qwen3.6 uncensored heretic

Related

Total size
28.3 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-31 22:05

Files by quantization

Auxiliary files 10 files 28.3 GB
model.safetensors 28.3 GB d01236ad download
tokenizer.json 19.1 MB 6f32ce20 download
config.json 14.9 KB e7a8b78e download
chat_template.jinja 11.5 KB 82faea87 download
README.md 5.21 KB f69e5a7c download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.24 KB a9eacca6 download
processor_config.json 1.16 KB 33818c7f download
generation_config.json 214 B 8a2e2eff download
recipe.yaml 207 B 3a29c3c6 download

README current version from Hugging Face


base_model: DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP
license: apache-2.0
library_name: transformers
tags:

  • fp8
  • compressed-tensors
  • quantized
  • vision-language
  • multimodal
  • qwen3.6
  • uncensored
  • heretic
    quantized_by: tacodevs

Qwen3.6-27B-Fable-Fusion-711 Uncensored-Heretic FP8 (vision-preserving)

Calibrated FP8 quantization of DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP that preserves vision-language capabilities, unlike vLLM's dynamic FP8 which destroys them on this architecture family.

TL;DR

  • ~30 GB FP8 (down from ~52 GB BF16), serves on a single 48 GB+ GPU with 16k context.
  • Vision verified intact after quantization (correctly reads fine image details — eye color, clothing trim, props — on real multimodal workloads).
  • Do NOT pass --quantization fp8 to vLLM — quantization config is baked in via compressed-tensors.
  • For structured/non-reasoning tasks, pass chat_template_kwargs: {"enable_thinking": false} — see Usage. Without it the model emits long tag-less "Thinking Process:" reasoning before answering.

About the source model

The source is DavidAU's multi-stage merge on the Qwen3.6-27B base (Qwen qwen3_5 architecture class: hybrid attention with Gated DeltaNet linear-attention layers + full attention every 4th layer, vision encoder included, 256k native context). Decensoring is via Heretic v1.2 with Arbitrary-Rank Ablation (ARA) — the source card measures 4/100 refusals vs 99/100 for the base model.

Note: the source repo ships a sidecar MTP (multi-token prediction) tensor file that is not part of the model index. It is not included in this quant; vLLM does not use it for standard serving.

Why dynamic FP8 destroys vision on this architecture

vLLM's runtime --quantization fp8 uses a single tensor-wide scale per Linear layer. The vision merger — the only bridge between the 1152-dim visual tower and the 5120-dim LM embedding space — has a much wider weight distribution than LM layers, so single-scale FP8 rounds its small-magnitude weights to zero. The LM then receives noise at image-token positions and silently hallucinates image descriptions from text alone. See the detailed write-up in the sibling repo tacodevs/Huihui-Qwen3.5-27B-Claude-4.6-Opus-abliterated-FP8.

What this checkpoint does differently

  • Per-channel weight scales (one scale per output channel, computed from actual weight distributions) instead of one global scale per layer.
  • Entire visual tower and merger kept in BF16 via the ignore list.
  • All Gated DeltaNet linear_attn modules kept in BF16 (hybrid-attention internals are not plain GEMMs and are excluded).
  • lm_head kept in BF16.
  • Dynamic per-token activation quantization at inference time.

Only the language model body (full-attention projections and MLP linears) is FP8.

Quantization details

  • Tool: llmcompressor (see recipe.yaml in this repo)
  • Scheme: FP8_DYNAMIC (per-channel weight scales, dynamic per-token activation scales)
  • Targets: all Linear layers
  • Excluded modules (207 total): re:.*visual.* (visual tower + merger), all linear_attn modules and norms, lm_head
  • Original size: ~52 GB BF16 → FP8 size: ~30 GB

Measured performance (RTX PRO 6000 Blackwell 96 GB, vLLM 0.26)

  • Decode: ~49 tok/s single-stream warm (identical to the BF16-recipe sibling Qwen3.5-27B FP8 on the same GPU).
  • TTFT ~0.8–1.0 s on multimodal requests (one 900px image + ~1k text tokens).
  • Vision: correct fine-grained image reading across all test runs (hair/eye color, clothing details, held objects, background).

Usage with vLLM

python -m vllm.entrypoints.openai.api_server \
    --model tacodevs/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-FP8 \
    --max-model-len 16384 \
    --gpu-memory-utilization 0.90 \
    --trust-remote-code

IMPORTANT: Do NOT pass --quantization fp8. The model already has its quantization config baked in via compressed-tensors; vLLM detects and uses the proper FP8 path automatically. Passing --quantization fp8 would re-quantize the already-FP8 weights and break everything.

Controlling thinking mode

The chat template supports Qwen's enable_thinking switch. By default the model produces extended reasoning without <think> tags (plain "Thinking Process:" markdown), which is easy to overrun token budgets with and hard to strip in streaming pipelines. For structured-output or latency-sensitive tasks, disable it per request:

{
  "model": "tacodevs/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-FP8",
  "chat_template_kwargs": {"enable_thinking": false},
  "messages": [...]
}

In our structured image-prompt workload this cut output from a truncated 500+ tokens to a complete 130–210 tokens and total latency from ~11 s to ~4.7 s, with format-perfect results.

Credits

  • Source merge: DavidAU
  • Decensoring: Heretic v1.2 (Arbitrary-Rank Ablation)
  • Base model: Qwen/Qwen3.6-27B
  • Quantization: tacodevs

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-31Add model card (quant recipe, vision-preservation rationale, enable_thinking ...a9c92f55.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration