← back to catalog · registered 2026-08-22 13:56

esatapedico/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-NVFP4-GGUF

esatapedico Qwen 27B GGUF second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/esatapedico%2FQwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-NVFP4-GGUF"
Response includes
  • classification m3
  • files 10
  • hub_downloads_all_time 6,930
  • author_summary 10 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
7K
1K last 30d - stable
Likes
2
Model age
2mo ago
created 2026-08-06
Downloads over time
Now7.5K→from2.4K↑217%
2.1K4.1K6K8K2.4K on Aug 127.5K on Oct 11AugSepOct
Aug 12 → Oct 11 · 49 snapshots · spans 60 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en multilingual
Tags
gguf nvfp4 qwen3.6 blackwell mtp speculative-decoding llama.cpp text-generation en multilingual base_model:DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP base_model:quantized:DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP

Related

Total size
64.3 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-19 08:20

Files by quantization

Auxiliary files 10 files 64.3 GB
Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-NVFP4-VERY-HIGH.gguf 18.3 GB fcb90677 download
Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-NVFP4-HIGH.gguf 16.3 GB 69ffaebc download
Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-NVFP4-MEDIUM.gguf 15.2 GB 71244344 download
Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-NVFP4-LOW.gguf 14.4 GB 2427a60d download
overrides.txt 22.0 KB edd0228e download
overrides-high.txt 22.0 KB 108d2568 download
overrides-medium.txt 22.0 KB 85334eb0 download
overrides-very-high.txt 22.0 KB 542d066f download
README.md 13.8 KB acf36630 download
.gitattributes 2.00 KB 9b3ebf32 download

README current version from Hugging Face


license: apache-2.0
base_model: DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP
pipeline_tag: text-generation
library_name: gguf
description: "Family of 4 tensor-spliced NVFP4 GGUFs of Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP (LOW / MEDIUM / HIGH / VERY-HIGH). Native NVFP4 backbone + per-tier extras (Q5_0/Q8_0/BF16 lm_head, IQ4_XS/Q6_K/BF16 embedding, MTP draft head). MTP speculative decoding. Blackwell sm_120."
tags:

  • gguf
  • nvfp4
  • qwen3.6
  • blackwell
  • mtp
  • speculative-decoding
  • llama.cpp
    language:
  • en
  • multilingual

Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-NVFP4-GGUF

A family of four tensor-spliced hybrid GGUFs of Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP:

  • Backbone (all 64 transformer blocks): native NVFP4, converted by us to GGUF from maci0's NVFP4 safetensors checkpoint.
  • Extras (output head, token embedding, MTP draft head): quantized per tier, sharing the same NVFP4 backbone across all four files.

The goal was simple: keep the NVFP4 prefill advantage, shrink the file below the 24 GB budget, and get faster decode via a higher-quality MTP draft head. On our dual 16 GB Blackwell setup all tiers fit and run. The MEDIUM/HIGH/VERY-HIGH tiers additionally ship with higher-precision output tensors to reduce the rare repetition-loop failure mode described below. See the notes before treating any numbers as meaningful.

Vision works too. The GGUFs themselves are text-only, but the model is vision-capable: pair it with the vision projector from DavidAU's NEO-MAX-MTP-GGUF repo, unmodified. DavidAU publishes mmproj-BF16.gguf, mmproj-F16.gguf, and mmproj-F32.gguf; we used the BF16 one ourselves (verified byte-identical), but any of the three works. No separate mmproj upload is needed here; just point --mmproj at one of his files.

Follow along & support

I post updates on new conversions, benchmarks, and what I'm working on over on Ko-fi. Follow along there to keep up with new releases and the work in progress. If you'd like to support more of it, a coffee is always welcome. I do this on consumer hardware and like seeing how far it goes. More is on the way.

☕ ko-fi.com/esatapedico. Updates, work-in-progress, and an optional coffee.

The four tiers

All tiers share the identical 496-tensor NVFP4 backbone (13.70 GB). They differ only in the 10 "extra" tensors (output.weight = LM head, token_embd.weight = embedding, and the 8 MTP draft-head tensors blk.64.*):

Tier File lm_head (output.weight) token_embd MTP head Size
LOW ...-NVFP4-LOW.gguf Q5_0 (0.87 GB) IQ4_XS IQ4_XS 15.49 GB
MEDIUM ...-NVFP4-MEDIUM.gguf Q8_0 (1.35 GB) Q6_K IQ4_XS 16.34 GB
HIGH ...-NVFP4-HIGH.gguf BF16 (2.54 GB) Q6_K IQ4_XS 17.53 GB
VERY-HIGH ...-NVFP4-VERY-HIGH.gguf BF16 BF16 BF16 19.65 GB

Tensor layout per tier (all 1,858 tensors; only the extra-tensor types differ):

GGML type Tensors Size Component
NVFP4 496 13.70 GB all 64 transformer blocks (attention QKV/output, FFN gate/up/down, gates); identical in all tiers
F32 1,352 0.01 GB norms, gates, scales
tier-dependent 10 see table output.weight, token_embd.weight, blk.64.* MTP draft block (incl. nextn.eh_proj)

The MTP draft head is embedded in the GGUF, so no separate drafter file is needed. Enable it in llama.cpp with --spec-type draft-mtp.

Why these extra tensors? The MTP draft head's job is to predict tokens the main model will accept. Keeping it at IQ4_XS instead of NVFP4 preserves draft quality, which is what makes speculative decoding pay off. The LM head precision is the main quality lever: Q5_0 (LOW) → Q8_0 (MEDIUM) → BF16 (HIGH/VERY-HIGH) progressively remove quantization noise from the exact tensor that determines the next-token distribution.

Anti-loop sampling

We previously recommended the DRY anti-loop sampler for the MEDIUM/HIGH/VERY-HIGH tiers; we no longer do — it interferes with verbatim reproduction of long strings (paths, identifiers, tool arguments), which matters for coding and tool use.

Attribution & provenance

This is a derivative work built entirely from existing Apache-2.0 artifacts. Nothing here was trained or fine-tuned. Credit belongs to:

  1. Alibaba / Qwen team for the base model, Qwen/Qwen3.6-27B (Apache-2.0): dense 27B, 64 layers, Gated DeltaNet + Gated Attention hybrid layout, native 262,144-token context.
  2. DavidAU for the fine-tune/merge DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP, and for the "LOW" quant design (special low-memory quants with Q5_0/Q6_K output tensor and IQ4_XS extras) from DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF.
  3. maci0 for the NVFP4 safetensors checkpoint maci0/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-NVFP4 (Apache-2.0). Note: maci0 ships safetensors; the GGUF conversion was done by us (see "How this was made").
  4. This repo's author for the tensor splice itself (combining the sources) and this write-up.

The NVFP4 tensors are native GGML type 40, preserved from the source checkpoint through GGUF conversion with no re-quantization and no dequant-to-requant round trip.

How this was made

  1. maci0's NVFP4 safetensors checkpoint was converted to GGUF by us using llama.cpp's convert_hf_to_gguf.py with --outtype auto (which preserves the native NVFP4 tensors as-is instead of dequantizing them). This produced the 19.65 GB all-NVFP4 backbone GGUF (the VERY-HIGH tier, unchanged).
  2. Each lower tier was produced with llama-quantize --tensor-type-file <overrides> from that parent, re-quantizing only the 10 extra tensors to the tier's chosen types and copying the NVFP4 backbone verbatim. The per-tier override maps are in this repo (overrides.txt, overrides-medium.txt, overrides-high.txt, overrides-very-high.txt).

No weights were retrained or re-quantized beyond the extra tensors; the NVFP4 backbone is byte-identical in all four files.

Repository contents

  • Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-NVFP4-LOW.gguf (15.49 GB)
  • Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-NVFP4-MEDIUM.gguf (16.34 GB)
  • Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-NVFP4-HIGH.gguf (17.53 GB)
  • Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-NVFP4-VERY-HIGH.gguf (19.65 GB)
  • overrides.txt (LOW), overrides-medium.txt, overrides-high.txt, overrides-very-high.txt: the per-tensor quantization-type maps from our conversion planning (770 entries each: every weight tensor and its target GGML type, e.g. blk.0.attn_qkv.weight=nvfp4, norms pinned to f32). Useful if you want to reproduce or audit the tensor layout without re-deriving it.

First observations (naive, single-run, not a benchmark)

We did not run a proper benchmark. What follows are informal first impressions from a handful of single-stream runs, included only so others know what to expect. Do not treat these as claims.

  • Hardware: 2x NVIDIA Blackwell 16 GB (RTX 5070 Ti + RTX 5060 Ti), split-mode: tensor, llama.cpp via LocalAI, flash attention on, quantized KV cache.
  • Prompt: one 180k-token payload (context 204,800). This payload was highly repetitive (it contained the same boilerplate unit 2,000 times), which likely inflates MTP acceptance and speedups; other prompts of different sizes behaved similarly, but this is far from comprehensive.
  • Sampling: temperature 0.6, top_p 0.95, top_k 20, min_p 0.
  • MTP: --spec-type draft-mtp, spec_n_max 6, spec_p_min 0.75.

The comparison rows:

  • maci0 NVFP4 row is our own GGUF conversion of maci0's safetensors checkpoint (the same --outtype auto conversion described above; now the VERY-HIGH tier), not a GGUF published by maci0.
  • IQ4XS LOW (DavidAU) row is DavidAU's published GGUF from his NEO-MAX-MTP repo.
  • NVFP4-LOW row is the LOW tier (same file that shipped in the previous version of this repo).
maci0 NVFP4 (= VERY-HIGH) IQ4XS LOW (DavidAU's GGUF) NVFP4-LOW
File size 19.65 GB 15.14 GB 15.49 GB
Prefill (first impression) 640.9 tok/s 563.5 tok/s 639.1 tok/s
Decode with MTP (first impression) 15.98 tok/s 19.78 tok/s ~23.5 tok/s
MTP acceptance (first impression) 0.880 0.878 ~0.92

Family comparison (naive, one run per tier, 180k payload to ~2k-token essay)

Each tier was measured with the same single 180k-token payload and a fresh LocalAI
process (pod restart) so only one model was resident at a time. All tiers share the
byte-identical NVFP4 backbone; only the 10 extra tensors differ.

Tier Prefill t/s Decode t/s MTP acceptance Mean draft len
LOW ~639 ~19.2 ~0.879 ~3.0
MEDIUM ~642 ~18.1 ~0.878 ~2.9
HIGH ~642 ~16.8 ~0.877 ~3.1
VERY-HIGH ~643 ~16.3 ~0.867 ~3.0

Naive first impressions, not a proper benchmark:

  • Prefill is effectively identical across tiers (~640 tok/s), which is expected:
    the 496-tensor NVFP4 backbone is byte-identical in all four files.
  • Decode rate decreases with tier size (LOW ~19.2, MEDIUM ~18.1, HIGH ~16.8,
    VERY-HIGH ~16.3 tok/s). The higher-precision lm_head/embedding tensors cost decode
    throughput (more bytes per token and slightly lower MTP acceptance).
  • No loops observed: every tier finished with finish_reason: stop and
    distinct-5-gram ratio >= 0.98 on the output. Repetition-analysis flags on the
    repetitive-payload runs traced back to the model quoting source sentences verbatim,
    not to degeneration.
  • Context used for these runs: the deployed defaults (LOW 307,200 / MEDIUM
    307,200, with 262,144 kept as the documented training context / HIGH 221,184 /
    VERY-HIGH 204,800). The larger
    weights of HIGH/VERY-HIGH do not fit a 307,200-token KV cache on a fresh process
    with the 16 GB per-GPU split, so those tiers run at lower contexts. All tiers
    stopped around 1.7k-2.9k output tokens despite the 20k request, so decode rates
    come from the actually-decoded window.

As with everything above: do not treat these as claims. Single-run, single-hardware,
repetitive payload.

Known caveat: rare non-deterministic repetition loop

During development, one out of five long-generation runs degenerated: the model repeated itself until hitting max_tokens exactly. The signatures were:

  • finish_reason: length at exactly the token cap (20,000)
  • decode rate far above the typical range seen in our other runs (~38 tok/s where ~20 to 24 was typical)
  • MTP acceptance creeping toward 0.98 (repetition is trivially draftable, so it inflates acceptance)

Four subsequent runs, including the exact same 180k to 20k scenario on a fresh process each time, did not reproduce it, and all finished with finish_reason: stop. The cause is not yet fully understood; treat it as a rare, non-deterministic failure mode to watch for, not a guaranteed one. If you see these three signatures together in a long generation, this is likely what is happening.

Mitigation: see "Anti-loop sampling" above. The tier differences remain in the model-side precision of the extra tensors.

Usage

llama.cpp / llama-server

llama-server \
  --model Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-NVFP4-MEDIUM.gguf \
  --mmproj <optional: Qwen3.6-27B mmproj-BF16.gguf> \
  --ctx-size 204800 \
  --flash-attn on \
  --spec-type draft-mtp \
  --spec-draft-n-max 6 \
  --spec-draft-p-min 0.75 \
  --temp 0.6 --top-p 0.95 --top-k 20
  • Requires a recent llama.cpp with NVFP4 (GGML type 40) CUDA kernels and sm_120 support (Blackwell).
  • Requires the draft-mtp spec path (merged upstream as LLAMA_CONTEXT_TYPE_MTP).
  • For vision input, point --mmproj at one of DavidAU's mmproj-BF16.gguf / mmproj-F16.gguf / mmproj-F32.gguf files from DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF (see the vision note above).
  • --spec-draft-p-min 0.75 matters: without it the drafter wastes steps on low-confidence tokens and the speedup shrinks.
  • MTP performance is hardware-dependent: try --spec-draft-n-max values 1 through 6 and keep whatever is fastest on your system.

Memory footprint

  • Weights: 15.49 / 16.34 / 17.53 / 19.65 GB per tier (all fit comfortably under 24 GB)
  • KV cache at 204,800 ctx with q4_0/q4_0: ~3 GB
  • Total on a 16+16 GB dual-GPU setup: ~18.5 to 22.7 GB depending on tier, leaving headroom

License

Apache-2.0, identical to every upstream artifact. The base model license governs; GGUF conversion and quantization are transformations, not new training. When redistributing, please retain attribution to Qwen (Alibaba), DavidAU, and maci0 as above.

"Qwen" is a trademark of Alibaba. Trademarks are used here only to identify upstream models; this repository is not affiliated with, sponsored by, or endorsed by Alibaba, DavidAU, or maci0.

Note on this card

This model card was written by an AI assistant at the request of the repository author, who did the engineering. As with any AI-generated text, there may be errors; please verify anything important (hashes, sizes, commands) against the file itself before relying on it.

README history 13 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-19Upload README.md with huggingface_hub194aab313.8 KB
    Loading...
  2. 2026-08-15Update model card: no longer recommend DRY anti-loop samplingfe3a6d113.4 KB
    Loading...
  3. 2026-08-12MEDIUM deployed context 307200 (training ctx 262144 documented)2db2eff14.2 KB
    Loading...
  4. 2026-08-12Note final deployed contexts (LOW 307200 / MEDIUM 262144 / HIGH 221184 / VERY...f1d209714.2 KB
    Loading...
  5. 2026-08-12Add family naive observations from 180k single-run tests84a2cf114.1 KB
    Loading...
  6. 2026-08-11DRY anti-loop sampler applies to all four tiers incl. LOW01d33a112.9 KB
    Loading...
  7. 2026-08-11Clean up card: remove em/en dashes4e45d9e12.9 KB
    Loading...
  8. 2026-08-11Add description to model card frontmatter0c3434f13 KB
    Loading...
  9. 2026-08-11Update model card to family wording (4 tiers)63505bd12.6 KB
    Loading...
  10. 2026-08-06Update README: full model name with MTP marker7e8631910.1 KB
    Loading...
  11. 2026-08-06Update vision note: any DavidAU mmproj variant works (BF16/F16/F32)6a615b710.1 KB
    Loading...
  12. 2026-08-06Update model card: add links to base models, remove YAML reference362316b9.9 KB
    Loading...
  13. 2026-08-06Add README.mda0bb3029.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration