← back to catalog · registered 2026-08-22 13:56

groxaxo/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-W8A8

groxaxo Qwen 24B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/groxaxo%2FQwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-W8A8"
Response includes
  • classification m3
  • files 22
  • benchmarks 11 entries
  • hub_downloads_all_time 29,303
  • author_summary 27 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
29K
58 last 30d - cooling
Likes
1
Model age
3mo ago
created 2026-06-16
Downloads over time
Now29.3K→from50↑58,558%
010.8K21.5K32.3K50 on Jun 1729.3K on Oct 1129.3K on Oct 7JunJulAugSepOct
Jun 17 → Oct 11 · 56 snapshots · spans 116 days

Benchmarks

Benchmark Score Source
Entertainment 1.7 UGI
Hazardous 2.4 UGI
Natural Intelligence 26.28 UGI
Political lean -26.8% UGI
Sensitive-Info 18.31 UGI
SocPol 1.6 UGI
UGI 43.88 UGI
Willingness (10) 9.5 UGI
W10-Adherence 10 UGI
W10-Direct 9 UGI
Writing 38.41 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
safetensors qwen3_5 auto-round compressed-tensors int8 w8a8 vllm mtp base_model:llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved base_model:quantized:llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved license:apache-2.0 8-bit

Related

Total size
29.1 GB
Files
22
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-22 08:06

Files by quantization

Auxiliary files 22 files 29.1 GB
model-00004-of-00010.safetensors 3.00 GB e9d8e058 download
model-00007-of-00010.safetensors 3.00 GB b2ee5149 download
model-00005-of-00010.safetensors 2.95 GB b31d7dd8 download
model-00001-of-00010.safetensors 2.94 GB 91dfae82 download
model-00002-of-00010.safetensors 2.92 GB 80de5362 download
model-00003-of-00010.safetensors 2.92 GB 398503c3 download
model-00006-of-00010.safetensors 2.92 GB 5705b185 download
model-00008-of-00010.safetensors 2.90 GB 6da47816 download
model-00009-of-00010.safetensors 2.37 GB 2e9b8089 download
model-00010-of-00010.safetensors 2.37 GB abed7b0a download
model_extra_tensors.safetensors 810 MB 713b0faf download
tokenizer.json 19.1 MB 6f32ce20 download
model.safetensors.index.json 160 KB d78e97c4 download
config.json 14.6 KB c6267e65 download
chat_template.jinja 11.5 KB 03a040a2 download
quantization_config.json 10.5 KB 0b12520f download
README.md 5.25 KB 2df25131 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.24 KB a9eacca6 download
processor_config.json 1.16 KB 33818c7f download
preprocessor_config.json 390 B 2ea84a43 download
generation_config.json 214 B 3f25ead4 download

README current version from Hugging Face


license: apache-2.0
base_model: llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved
tags:

  • auto-round
  • compressed-tensors
  • int8
  • w8a8
  • vllm
  • qwen3_5
  • mtp

Qwen3.6-27B-uncensored-heretic-v2 — W8A8 (AutoRound INT8 Dynamic, MTP preserved)

Overview

Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-W8A8 is an 8-bit checkpoint designed to reduce inference memory requirements, published by groxaxo.
It is intended for open-source evaluation, reproducible experimentation, and compatible local or
hosted inference workflows. The wording below is deliberately limited to what can be verified
from this repository's metadata and artifacts.

The repository name identifies a behavior-modified or reduced-filtering lineage. That label describes the source or conversion history; it is not a guarantee of unrestricted behavior in every prompt or runtime. Test outputs carefully before sharing or deploying them.

At a glance

Field Details
Format INT8
Source / base llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved
Intended task image-text-to-text
License apache-2.0

What is included

  • *.safetensors (11 files)
  • config.json
  • generation_config.json
  • tokenizer.json
  • tokenizer_config.json
  • processor_config.json
  • chat_template.jinja
  • quantization_config.json
  • Additional configuration, tokenizer, processor, or shard files (20 visible artifacts total)

Quick start

Runtime selection

Load the checkpoint with a runtime that supports its architecture and 8-bit weight format. Keep
the repository's configuration and tokenizer/processor files together with the weight files.

Compatibility and responsible use

  • Use a runtime that explicitly supports this format, architecture, and modality.
  • Keep configuration, tokenizer, processor, projection, and weight files from the same revision together.
  • Review the source model card and license before redistribution or deployment.
  • Hardware needs depend on parameter count, context length, cache precision, quantization, and concurrency.
  • Report reproducible issues with the runtime version, hardware, launch command, and a minimal example.

Quantization or conversion changes numerical behavior, memory use, and throughput relative to the source checkpoint; validate quality on your own workload.

Generated outputs may be inaccurate or unsuitable for a given use case. Users are responsible for
testing behavior, applying appropriate safeguards, and complying with applicable licenses and laws.

INT8 W8A8 dynamic (per-token activation) quantization of
llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved,
produced with AutoRound 0.13.1 and exported in the
compressed-tensors (auto_round:llm_compressor) format for vLLM.

Same recipe as the Darwin-28B W8A8 build.

Quantization recipe

  • Scheme: INT8 — channel-wise INT8 weights, INT8 dynamic per-token activations (CUTLASS INT8 path in vLLM).
  • Algorithm: RTN (round-to-nearest, data-free). INT8 + dynamic activations are near-lossless.
  • Quantized: language-model Linear weights only (DeltaNet in_proj_*/out_proj + full-attn q/k/v/o_proj + MLP gate/up/down).
  • Preserved in BF16: native MTP module (mtp.*, via --ignore_layers mtp), the vision tower (model.visual.*, auto-skipped by AutoRound's text-module-only VLM path), lm_head, embed_tokens, all norms.
  • Hardware: 2 × RTX 3090.

Architecture: Qwen3.5-generation hybrid (qwen3_5, 64 layers, 3:1 DeltaNet/full-attention, hidden 5120, vocab 248320), multimodal, with a 1-layer native MTP head for speculative decoding.

Serving (vLLM, TP2) — recommended: dense (no MTP)

vllm serve groxaxo/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-W8A8 \
  --tensor-parallel-size 2 \
  --max-model-len 262144 \
  --kv-cache-dtype fp8 \
  --reasoning-parser qwen3 \
  --trust-remote-code

MTP note (measured)

The native MTP head is preserved in BF16, but on this finetune it is non-predictive: measured
~0% draft-token acceptance (2 of 6254), so enabling speculative decoding is actually slower
(~22.5 tok/s with MTP vs ~27.7 tok/s dense, single-stream on 2×RTX 3090). This was isolated to the
checkpoint's MTP head, not the quantization or server:

  • Same vLLM + method:mtp on the standard Qwen3.6-27B (AEON FP8) gives 78% acceptance / ~2.1× speedup — so the setup is correct.
  • 0% acceptance persists with kv-cache-dtype=auto (not a KV-quant artifact).
  • The W8A8 quant is near-lossless (KL≈0.02 nats/token) and preserves the MTP head structure intact.

The heretic-v2 finetune appears to have diverged its main weights from the base MTP head, so the head's
drafts no longer match. Serve dense for best speed. The MTP weights remain in this repo for anyone
who retrains/distills the head against this finetune. The vision tower is intact (drop
--hf-overrides '{"language_model_only": true}' to use it; add it to save VRAM for text-only).

README history 7 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-22Polish model card overview and usage notes2e9c57e5.3 KB
    Loading...
  2. 2026-08-22Polish model card overview and usage notesbee78615.4 KB
    Loading...
  3. 2026-08-22Polish model card overview and usage notes1487b585.3 KB
    Loading...
  4. 2026-08-22Polish model card overview and usage notes929ec425.3 KB
    Loading...
  5. 2026-08-22Polish model card overview and usage notes6eed8144.5 KB
    Loading...
  6. 2026-06-17Upload README.md with huggingface_hubef216232.8 KB
    Loading...
  7. 2026-06-16Upload folder using huggingface_hub3567e6c2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration