← back to catalog · registered 2026-08-22 13:56

huginnfork/Qwen3.6-27B-uncensored-heretic-v2-mtp-NVFP4A16

huginnfork Qwen 9.4B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/huginnfork%2FQwen3.6-27B-uncensored-heretic-v2-mtp-NVFP4A16"
Response includes
  • classification m3
  • files 19
  • hub_downloads_all_time 1,626
  • author_summary 4 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
2K
69 last 30d - cooling
Likes
2
Model age
5mo ago
created 2026-04-26
Downloads over time
Now1.7K→from469↑253%
4108641.3K1.8K469 on Apr 291.7K on Oct 11AprMayJunJulAugSepOct
Apr 29 → Oct 11 · 63 snapshots · spans 165 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
safetensors qwen3_5 nvfp4 nvfp4a16 compressed-tensors multimodal mtp speculative-decoding heretic abliteration image-text-to-text conversational

Related

Total size
26.6 GB
Files
19
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-27 09:47

Files by quantization

Auxiliary files 19 files 26.6 GB
model-00001-of-00006.safetensors 5.00 GB 13b5b7a1 download
model-00005-of-00006.safetensors 5.00 GB 6e71199d download
model-00002-of-00006.safetensors 5.00 GB 11bf8c73 download
model-00004-of-00006.safetensors 4.95 GB d68abb5a download
model-00003-of-00006.safetensors 4.94 GB 275479c4 download
model-00006-of-00006.safetensors 1.70 GB 36d8252f download
tokenizer.json 19.1 MB f399b3cd download
model.safetensors.index.json 164 KB 9af6ea5b download
config.json 24.1 KB 2a406c5e download
chat_template.jinja 7.67 KB 3bba6834 download
README.md 3.74 KB 07a0fece download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.20 KB 920f0987 download
recipe.yaml 593 B 02d34e5e download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
kld_heretic_a16_v2.json 303 B 4a859f9d download
kld_heretic_a16_vs_base_v2.json 301 B 7e9e87c1 download
generation_config.json 213 B e5597d6c download

README current version from Hugging Face


license: apache-2.0
base_model:

  • huginnfork/Qwen3.6-27B-uncensored-heretic-v2-mtp
  • Qwen/Qwen3.6-27B
    pipeline_tag: image-text-to-text
    tags:
  • nvfp4
  • nvfp4a16
  • compressed-tensors
  • multimodal
  • mtp
  • speculative-decoding
  • heretic
  • abliteration
    base_model_relation: quantized

huginnfork/Qwen3.6-27B-uncensored-heretic-v2-mtp-NVFP4A16

NVFP4A16 quantisation derived from the heretic-abliterated llmfan46/Qwen3.6-27B-uncensored-heretic-v2 of Qwen/Qwen3.6-27B, with the MTP head and vision tower preserved in bf16.

Provenance

  • Base: Qwen/Qwen3.6-27B (bf16)
  • Abliteration tool: heretic v1.2.0 (ARA) by Philipp Emanuel Weidmann (p-e-w)
  • Abliterated source weights: llmfan46/Qwen3.6-27B-uncensored-heretic-v2 — the heretic-derived bf16 abliteration of Qwen/Qwen3.6-27B
  • MTP head: re-grafted from Qwen/Qwen3.6-27B (15 tensors, ~810 MB bf16) so SGLang/vLLM speculative decoding (--speculative-algo NEXTN) works
  • Vision tower: model.visual.* preserved in bf16 (333 tensors)

Quantisation

  • Format: NVFP4A16 W4A16 (4-bit FP4 weights with FP8 E4M3 per-group scales, group_size=16; bf16 activations) via llm-compressor 0.10 + compressed-tensors 0.14
  • Scheme: see recipe.yaml — NVFP4A16 keeps activations in bf16, which is the dominant KLD source under W4A4 on this hybrid stack.
  • Kept in bf16 (quantization_config.ignore): lm_head, all model.visual.*, all mtp.*, and the entire linear_attn Mamba/SSM block (in_proj_*, out_proj, conv1d)

KL divergence measurements

KLD computed with eval_kld.py — per-token KLD averaged over 8 samples from neuralmagic/calibration (LLM split), max_seq=1024. Max sample KLD is the highest single-sample mean (catches outliers that the overall mean hides).

Comparison Mean KLD (nats) Max sample KLD Samples max_seq
vs Qwen3.6-27B base 0.1012 0.2213 8 1024

Note: this pipeline always uploads the resulting checkpoint. Consult the KL
divergence numbers above to judge whether the result is acceptable for your
use case.

Perplexity (wikitext-2-raw)

Wikitext-2-raw test split, non-overlapping chunks of 2048 tokens, computed with eval_ppl.py. Same tokenizer for every row so the numbers compare apples-to-apples.

Model Perplexity Tokens scored Dataset seq
Qwen3.6-27B base (bf16) 7.3057 296907 wikitext/wikitext-2-raw-v1/test 2048
heretic-v2-mtp (bf16) 7.4619 296907 wikitext/wikitext-2-raw-v1/test 2048
this checkpoint 7.8022 296907 wikitext/wikitext-2-raw-v1/test 2048

Inference

transformers (text + vision; MTP not exercised)

from transformers import AutoModelForImageTextToText, AutoProcessor
import torch

repo = "huginnfork/Qwen3.6-27B-uncensored-heretic-v2-mtp-NVFP4A16"
proc = AutoProcessor.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForImageTextToText.from_pretrained(
    repo, dtype=torch.bfloat16, device_map="auto", trust_remote_code=True,
)

vLLM (NVFP4A16 + MTP speculative decoding)

vllm serve huginnfork/Qwen3.6-27B-uncensored-heretic-v2-mtp-NVFP4A16 \
    --trust-remote-code \
    --gpu-memory-utilization 0.85 \
    --max-model-len 8192 \
    --quantization compressed-tensors \
    --speculative-config '{"method":"qwen3_5_mtp","num_speculative_tokens":1}'

FP4 weights are unpacked to bf16 at compute time. Native FP4 GEMM requires Blackwell (SM100+); on older GPUs vLLM dequantises at runtime.

README history 10 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-27Refresh KLD numbers + add wikitext-2-raw perplexity tablee05fc9d3.7 KB
    Loading...
  2. 2026-04-27Rename huginnfork/Qwen3.6-27B-heretic-NVFP4A16 -> huginnfork/Qwen3.6-27B-unce...12edb273.1 KB
    Loading...
  3. 2026-04-27Fix doubled curly braces in vLLM speculative-config JSON64d2a003.1 KB
    Loading...
  4. 2026-04-27Set base_model + base_model_relation metadata9cbf42d3.1 KB
    Loading...
  5. 2026-04-26Update README2e16eb13.1 KB
    Loading...
  6. 2026-04-26Initial upload140015e3.1 KB
    Loading...
  7. 2026-04-26Update README90a7f7d3.1 KB
    Loading...
  8. 2026-04-26Update README15f21073 KB
    Loading...
  9. 2026-04-26Initial upload4a173183 KB
    Loading...
  10. 2026-04-26Initial uploaded577912.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration