← back to catalog · registered 2026-08-22 13:56

pocharlies/Qwen3.6-27B-uncensored-heretic-v2-NVFP4-lmheadW4

pocharlies Qwen 13B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/pocharlies%2FQwen3.6-27B-uncensored-heretic-v2-NVFP4-lmheadW4"
Response includes
  • classification m3
  • files 11
  • hub_downloads_all_time 274
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
274
51 last 30d - stable
Likes
0
Model age
2mo ago
created 2026-07-24
Downloads over time
Now300→from38↑689%
2512522632638 on Jul 22300 on Oct 11JulAugSepOct
Jul 22 → Oct 11 · 52 snapshots · spans 81 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
safetensors qwen3_5 nvfp4 vllm gb10 dgx-spark sm121 mtp speculative-decoding uncensored qwen3.6 text-generation

Related

Total size
17.4 GB
Files
11
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-24 20:00

Files by quantization

Auxiliary files 11 files 17.5 GB
model.safetensors 17.4 GB 17dc28f3 download
tokenizer.json 19.1 MB 6f32ce20 download
chat_template.jinja 11.5 KB 03a040a2 download
config.json 6.97 KB e1bfcf85 download
README.md 3.11 KB 3bdcc920 download
hf_quant_config.json 3.02 KB 9d73e482 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.13 KB c487bad4 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 226 B 2dd033e0 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-NVFP4
  • Qwen/Qwen3.6-27B
    language:
  • en
    pipeline_tag: text-generation
    tags:
  • nvfp4
  • vllm
  • gb10
  • dgx-spark
  • sm121
  • mtp
  • speculative-decoding
  • uncensored
  • qwen3.6

Qwen3.6-27B-uncensored-heretic-v2-NVFP4-lmheadW4

What this is: the llmfan46 heretic-v2 NVFP4 uncensored checkpoint with a single, high-impact optimization: the lm_head re-quantized from BF16 to NVFP4 W4A16.

Community uncensored/abliterated checkpoints almost always ship the lm_head in BF16 (here ~2.5 GB over a 248,320-token vocabulary). That tensor is read every decode step, so on a memory-bandwidth-bound GPU it dominates the per-token cost and makes decode slow. NVIDIA's own NVFP4 checkpoints quantize the lm_head to 4-bit — which is why they decode fast. We simply brought this checkpoint in line.

Result (measured on NVIDIA GB10 / DGX Spark, sm_121)

metric heretic-v2 (BF16 lm_head) this (NVFP4 lm_head)
decode tok/s @ 8K (single stream) ~28 41.2
decode tok/s @ 16K — 40.6
decode tok/s @ 32K — 39.5
prefill tok/s @ 8K ~2100 2155

+~47% decode, no quality regression (needle-in-haystack 8K/32K HIT, coherence preserved), uncensored behaviour preserved. On our cluster it now decodes faster than the dense Qwen3.6-27B (~35 tok/s).

How it was made

Standalone lm_head quantization (loads only the lm_head tensor, ~21 GB peak — avoids OOM on the 120 GB unified memory) with NVIDIA modelopt NVFP4QTensor.quantize(W, block_size=16), producing the modelopt tensor layout (weight packed uint8 [V, H/2], weight_scale FP8-E4M3 [V, H/16], weight_scale_2 FP32, input_scale FP32). The format is verified bit-for-bit against a reference NVFP4 lm_head before the shard is rewritten, and lm_head is removed from quantization_config.ignore. Everything else (MoE/attention weights, native MTP tensors) is untouched. Script: see the companion repo.

Serving (vLLM on GB10, custom sm121 build)

vllm serve pocharlies/Qwen3.6-27B-uncensored-heretic-v2-NVFP4-lmheadW4 \
  --served-model-name qwen36-27b-uncensored-nvfp4 \
  --max-model-len 229376 \
  --kv-cache-dtype fp8 \
  --quantization modelopt \
  --attention-backend flashinfer \
  --max-num-seqs 16 --max-num-batched-tokens 32768 \
  --gpu-memory-utilization 0.60 \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
  --limit-mm-per-prompt '{"image":64}' \
  --trust-remote-code

KV cache nvfp4 is NOT possible on GB10 (FlashInfer requires sm100f; GB10 is sm121) → use --kv-cache-dtype fp8.

Attribution & license

Apache-2.0. Derivative of llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-NVFP4, itself derived from Qwen/Qwen3.6-27B. Only the lm_head tensor was re-quantized; all other weights are unchanged. Abliteration/uncensoring was performed upstream by llmfan46 — this repo adds only a quantization/performance optimization.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-24Add model card91bf99b3.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration