← back to catalog · registered 2026-08-22 13:56

huginnfork/Qwen3.6-27B-uncensored-heretic-v2-mtp-FP8

huginnfork Qwen 17B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/huginnfork%2FQwen3.6-27B-uncensored-heretic-v2-mtp-FP8"
Response includes
  • classification m3
  • files 22
  • hub_downloads_all_time 8,480
  • author_summary 4 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
8K
69 last 30d - cooling
Likes
5
Model age
5mo ago
created 2026-04-26
Downloads over time
Now8.5K→from246↑3,358%
03.1K6.2K9.3K246 on Apr 298.5K on Oct 118.5K on Oct 9AprMayJunJulAugSepOct
Apr 29 → Oct 11 · 63 snapshots · spans 165 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 95 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Tags
safetensors qwen3_5 fp8 compressed-tensors mtp speculative-decoding multimodal image-text-to-text conversational base_model:huginnfork/Qwen3.6-27B-uncensored-heretic-v2-mtp base_model:quantized:huginnfork/Qwen3.6-27B-uncensored-heretic-v2-mtp license:apache-2.0

Related

Total size
35.8 GB
Files
22
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-30 11:02

Files by quantization

Auxiliary files 22 files 35.8 GB
model-00007-of-00008.safetensors 5.00 GB e39c9663 download
model-00006-of-00008.safetensors 4.99 GB 86174bbe download
model-00003-of-00008.safetensors 4.97 GB 14258bb5 download
model-00005-of-00008.safetensors 4.97 GB d329aac0 download
model-00004-of-00008.safetensors 4.97 GB 5887cfc9 download
model-00001-of-00008.safetensors 4.95 GB de78fd4f download
model-00002-of-00008.safetensors 4.94 GB 81e6531c download
model-00008-of-00008.safetensors 1.02 GB 14d1340c download
tokenizer.json 19.1 MB 225fe96e download
model.safetensors.index.json 128 KB 67e1f5bd download
config.json 33.0 KB 3f9fa11c download
chat_template.jinja 7.82 KB 09c96b90 download
README.md 2.02 KB 604e4f67 download
recipe.yaml 1.83 KB 6e73aa86 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.20 KB 920f0987 download
kld_heretic_attnbf16.json 662 B b7969abf download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
kld_fp8_vs_heretic.json 306 B 88f73f9d download
kld_fp8_vs_base.json 304 B a5b22328 download
generation_config.json 226 B 16d319af download

README current version from Hugging Face


base_model:

  • huginnfork/Qwen3.6-27B-uncensored-heretic-v2-mtp
    base_model_relation: quantized
    license: apache-2.0
    pipeline_tag: image-text-to-text
    tags:
  • qwen3_5
  • fp8
  • compressed-tensors
  • mtp
  • speculative-decoding
  • multimodal

Qwen3.6-27B-uncensored-heretic-v2-mtp-FP8

FP8 (compressed-tensors, FP8_DYNAMIC W8A8) quantisation of huginnfork/Qwen3.6-27B-uncensored-heretic-v2-mtp,
attnbf16 variant: the entire self-attention path is kept in bf16 and only the MLPs are FP8. Ships a
working MTP self-speculative-decoding head.

What's kept in bf16

lm_head, the MTP head, the vision tower, the whole linear_attn (Gated-DeltaNet / SSM) block,
and the entire self_attn path (incl. the attention output gate fused into q_proj on Qwen3.5/3.6).
Only the ~17 B MLP params are FP8. This keeps quantisation off the multiplicative attention gate and the
16 long-range full-attention layers, at a cost of ~+1.5 GiB (~4.6 %) vs a fully-quantised FP8 build. See
recipe.yaml.

Accuracy vs the bf16 parent: KLD ≈ 0.0203 nats (kld_heretic_attnbf16.json), measured per-token on
neuralmagic/calibration (8 samples, seq 1024). This is markedly lower than a plain-attention FP8 build
of the same base.

Speculative decoding (MTP)

The bf16 MTP head is declared in quantization_config.ignore so vLLM loads it correctly. (A bf16 MTP head
regrafted into a compressed-tensors quant is otherwise mis-loaded and yields 0 % draft acceptance — this
build fixes that.)

Measured MTP acceptance: 73.0 % (vLLM 0.26.0, Blackwell, greedy,
--speculative-config '{"method":"mtp","num_speculative_tokens":3}').

vllm serve huginnfork/Qwen3.6-27B-uncensored-heretic-v2-mtp-FP8 \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
  --max-num-seqs 32

--max-num-seqs 32 (or lower) is required — Qwen3.6 is a hybrid linear-attention model whose Mamba cache
otherwise runs out of blocks at the default max_num_seqs.

README history 7 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-30rebuild: attnbf16 FP8 + working MTP head8982bb32 KB
    Loading...
  2. 2026-04-27Refresh KLD numbers + add wikitext-2-raw perplexity table60ac10d3.7 KB
    Loading...
  3. 2026-04-27Rename huginnfork/Qwen3.6-27B-heretic-FP8 -> huginnfork/Qwen3.6-27B-uncensore...b5d25613 KB
    Loading...
  4. 2026-04-27Fix doubled curly braces in vLLM speculative-config JSONd8411073 KB
    Loading...
  5. 2026-04-27Set base_model + base_model_relation metadata5f0824c3 KB
    Loading...
  6. 2026-04-26Update README0d6c1003 KB
    Loading...
  7. 2026-04-26Initial upload6d883f33 KB
    Loading...

Discussions 1 thread

  1. 2026-06-12It would be great if you could quantize FP8 again using this update version of …open2 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration