← back to catalog · registered 2026-08-22 13:56

maci0/Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking-NVFP4

maci0 Qwen 37B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/maci0%2FQwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking-NVFP4"
Response includes
  • classification m3
  • files 17
  • hub_downloads_all_time 13,889
  • author_summary 8 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
14K
703 last 30d - cooling
Likes
12
Model age
3mo ago
created 2026-06-18

Training datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now14K→from444↑3,051%
05.1K10.2K15.3K444 on Jun 1714K on Oct 11JunJulAugSepOct
Jun 17 → Oct 11 · 57 snapshots · spans 116 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_5 image-text-to-text nvfp4 fp4 w4a4 gptq quantized compressed-tensors llm-compressor vllm

Related

Total size
24.7 GB
Files
17
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-04 01:51

Files by quantization

Auxiliary files 17 files 24.8 GB
model.safetensors 23.9 GB 60f8b746 download
model-towers.safetensors 879 MB c4d6b80a download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 339 KB a0309392 download
nvfp4_benchmarks.png 65.3 KB b492cc8c download
README.md 18.1 KB eaa5e82c download
config.json 11.4 KB c87674ec download
chat_template.jinja 7.58 KB a8755d82 download
QUANTIZATION.md 7.04 KB 54256b97 download
BENCHMARKS.md 2.28 KB bc3fdc1a download
.gitattributes 1.59 KB 88b36108 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.14 KB 1d134cd2 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
recipe.yaml 326 B 0ed8315b download
generation_config.json 214 B 3f25ead4 download

README current version from Hugging Face


base_model: DavidAU/Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking
base_model_relation: quantized
license: apache-2.0
library_name: transformers
pipeline_tag: image-text-to-text
language:

  • en
  • zh
    tags:
  • nvfp4
  • fp4
  • w4a4
  • gptq
  • quantized
  • compressed-tensors
  • llm-compressor
  • vllm
  • qwen3_5
  • vision-language
  • thinking
  • uncensored
  • creative
    datasets:
  • TeichAI/claude-4.5-opus-high-reasoning-250x
  • HuggingFaceH4/ultrachat_200k
  • m-a-p/Code-Feedback

RQ-40B-DKDUncensored
Qwen3.6-40B Deckard · NVFP4
40B dense VL · Claude 4.6 Opus reasoning distill · thinking · uncensored.
Params40B
Active40B (dense)
Size25.6 GB
Perplexity6.89
Refusalsn/a
Context256K
MTP headn/a

TL;DR: Qwen3.6-40B Deckard, quantized to NVFP4 (W4A4) for vLLM on NVIDIA Blackwell. 25.6 GB, wikitext-2 PPL 6.89, flagship reasoner, uncensored.

Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking NVFP4

NVFP4 (W4A4) quantization of
DavidAU/Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking,
packed in the compressed-tensors nvfp4-pack-quantized format with
llm-compressor. Weights are
quantized with GPTQ (error-compensated rounding) and an MSE observer, using a
domain-matched calibration blend.

Near-lossless versus bf16. SWE-bench Lite resolves land within one instance of the
bf16 source at every size, the 27B lm-eval gap averages under 1.5 points, and
wikitext-2 perplexity for this build is 6.89. See Benchmarks.

  • About 25.6 GB on disk versus about 73.7 GB for the bf16 source (about 35%).
  • Built for vLLM on NVIDIA Blackwell, where both the 4-bit weight and 4-bit
    activation paths are accelerated. On pre-Blackwell GPUs vLLM runs it weight-only.
  • Loading and generation verified in vLLM on an NVIDIA GB10 (Blackwell, sm_121).

Uncensored model. This is a quantization of an uncensored / abliterated
derivative. It follows instructions without content guardrails, including NSFW.
Behaviour and alignment are inherited entirely from the base model.

Benchmarks

benchmarks

Near-lossless versus the bf16 source:

  • SWE-bench Lite (agentic, mini-swe-agent, instances 0:20): resolves land within
    one instance of bf16 at every size (NVFP4 15/13, bf16 16/14 of 20).
  • lm-eval (27B pair, the clean apples-to-apples): average accuracy gap under 1.5
    points.
  • wikitext-2 perplexity (this 40B build, vLLM prompt-logprobs): 6.89.

Full head-to-head tables and method in BENCHMARKS.md.

Fidelity

Near-lossless versus the bf16 source, 25.6 GB vs 73.7 GB bf16 (~35%), at wikitext-2 perplexity 6.89. See Benchmarks for the full head-to-head. GPTQ error compensation and an MSE observer keep the drop from bf16 minimal; the header lists the full characteristics and Quantization covers the recipe.

Quickstart

Offline (vLLM)

NVFP4 activation acceleration needs a Blackwell-class GPU. The if __name__ == "__main__" guard is required for offline LLM(...) because the vLLM v1 engine
spawns workers.

from vllm import LLM, SamplingParams

def main():
    llm = LLM(
        model="maci0/Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking-NVFP4",
        max_model_len=16384,
    )
    msgs = [{"role": "user", "content": "Write the opening paragraph of a noir short story."}]
    sp = SamplingParams(temperature=1.0, top_p=0.95, top_k=20, max_tokens=2048)
    out = llm.chat(msgs, sp)
    print(out[0].outputs[0].text)

if __name__ == "__main__":
    main()

Server (OpenAI-compatible)

Recommended baseline for a single Blackwell GPU. The NVFP4 quantization is
auto-detected from config.json (compressed-tensors), so no quantization flag is
needed. --reasoning-parser qwen3 splits the <think> block into a separate
reasoning_content field.

vllm serve maci0/Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking-NVFP4 \
  --served-model-name qwen3.6-40b-nvfp4 \
  --tensor-parallel-size 1 \
  --max-model-len 131072 \
  --gpu-memory-utilization 0.90 \
  --kv-cache-dtype fp8 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice --tool-call-parser qwen3_coder

These parser flags are not auto-detected; you must pass them explicitly. Drop the
last line if you do not need tool calling; --enable-auto-tool-choice requires
--tool-call-parser.

Verified flag values for this model on vLLM:

Goal Add
Tool / function calling --enable-auto-tool-choice --tool-call-parser qwen3_coder
Text-only (skip vision tower, free KV cache) --language-model-only
Bound multimodal inputs --limit-mm-per-prompt '{"image":4,"video":1}'
Hour-scale video --media-io-kwargs '{"video":{"num_frames":-1}}' (and raise longest_edge in video_preprocessor_config.json)

Context notes:

  • The model supports up to 262144 tokens. Upstream guidance is to keep at least
    128K to preserve thinking quality, so --max-model-len 131072 is the recommended
    default. Go to 262144 if memory allows, or lower it if you hit OOM.
  • On unified-memory parts (e.g. GB10), --gpu-memory-utilization carves from RAM
    shared with the rest of the system. Use about 0.90 when this is the only model,
    and leave more headroom (about 0.80) when co-hosting other processes.

Python (OpenAI client)

from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="x")
r = client.chat.completions.create(
    model="qwen3.6-40b-nvfp4",
    messages=[{"role": "user", "content": "Write the opening paragraph of a noir short story."}],
)
print(r.choices[0].message.content)

curl

curl http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" -d '{
  "model": "qwen3.6-40b-nvfp4",
  "messages": [{"role": "user", "content": "Write the opening paragraph of a noir short story."}]
}'

KV-cache quantization

The NVFP4 here quantizes weights and activations, not the KV cache (the checkpoint
ships kv_cache_scheme: null). KV-cache quantization is a separate runtime vLLM
option. Only the 24 full-attention layers hold a standard KV cache; the gated
delta-net linear-attention layers use recurrent state and are unaffected.

  • FP8 KV cache (recommended, safe): --kv-cache-dtype fp8 (or fp8_e4m3). About
    2x KV savings with small quality cost. Used in the baseline command above.
  • TurboQuant (vLLM, experimental here): lower-bit KV quant via Hadamard
    rotation plus per-coordinate Lloyd-Max scalar quantization. Values:
    turboquant_k8v4 (FP8/4-bit, 2.6x, +1.17% PPL), turboquant_4bit_nc (3.8x),
    turboquant_k3v4_nc (~3.5x), turboquant_3bit_nc (4.9x). turboquant_k8v4 is
    the quality sweet spot. Caveat: TurboQuant uses a dedicated attention backend
    whose interaction with this model's linear-attention layers was not verified.
    Treat as experimental; prefer fp8 for a known-good KV quant.

Performance / backend notes (verified on vLLM)

  • FlashInfer is bundled and autotune is on by default. It is used automatically for
    the full-attention and NVFP4 GEMM paths on Blackwell; there is nothing to enable.
    Optional: VLLM_USE_FLASHINFER_SAMPLER=1 for faster sampling.
  • NVFP4 GEMM auto-selects cutlass FP4 on Blackwell. Do not set
    VLLM_NVFP4_GEMM_BACKEND (deprecated in 0.23.0). Leave
    VLLM_USE_NVFP4_CT_EMULATIONS=0 (the default; emulation is for pre-Blackwell).
  • Attention backend: leave on auto. This is a hybrid model, so vLLM assigns the
    per-layer backends (GDNAttentionBackend / LinearAttentionBackend)
    automatically. Forcing a single global attention backend breaks the
    linear-attention layers.
  • No sparse-attention knob applies. The efficiency comes from the hybrid 3:1
    linear:full attention layout, handled automatically.

About the base model

A 40B dense (not MoE) vision-language model expanded from Qwen3.6-27B, made
uncensored via Heretic, trained on the internal Deckard/PKD datasets (character,
depth, point of view) and on a Claude 4.6 Opus high-reasoning distillation set to
sharpen and stabilize reasoning.

  • 96 decoder layers: hybrid gated delta-net linear attention (72) plus full
    attention (24), dense MLP, plus a vision tower for image and video input.
  • 256K context (max_position_embeddings 262144).
  • Thinking mode by default (variable-length reasoning), with an instruct toggle.

Quantization

Scheme NVFP4, W4A4
Weight rounding GPTQ (Hessian-based error compensation), MSE observer
Weights FP4 (E2M1), group_size=16, tensor_group, symmetric, FP8 (E4M3) group scales
Activations FP4, dynamic per-group (dynamic: local), FP8 (E4M3) scales
Targets all language-model Linear layers, 744 modules (360 linear-attn projections + 288 MLP + 96 full-attn)
Kept in bf16 vision tower (model.visual.*), lm_head
Untouched gated delta-net Conv1d and SSM params (A_log, dt_bias), not Linear, never targeted

GPTQ is a quantization-time cost only. The output is the same
nvfp4-pack-quantized format with identical inference speed; GPTQ just chooses
better 4-bit values than plain round-to-nearest.

Calibration

512 samples, domain-matched to the model's actual traffic, max_seq_len=2048,
text-only path through the VL model:

source samples domain
TeichAI/claude-4.5-opus-high-reasoning-250x 250 long reasoning (the base model's own training data)
HuggingFaceH4/ultrachat_200k 150 general chat
m-a-p/Code-Feedback 112 code

Quality

GPTQ with the domain-matched calibration measurably beats plain round-to-nearest. The
fused layers (q/k/v, gate/up) share one NVFP4 global scale, so vLLM does not warn or
fall back. Measured wikitext-2 perplexity for this build is 6.89 (see
Benchmarks); the gain is expected to be larger on the model's own
domains (reasoning, creative, code), which wikitext does not cover.

Recommended sampling

Thinking mode is the default.

  • Thinking, general: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, repetition_penalty=1.0
  • Thinking, precise coding: temperature=0.6, top_p=0.95, top_k=20
  • Instruct / non-thinking: temperature=0.7, top_p=0.80, top_k=20, presence_penalty≈1.5
  • If the model loops on thin prompts, add a one-line system prompt (e.g. Be vivid and precise.) and/or set repetition_penalty 1.05 to 1.1.

To run instruct (non-thinking), set {%- set enable_thinking = false %} in the
Jinja chat template, or pass
extra_body={"chat_template_kwargs": {"enable_thinking": false}} on OpenAI-compatible
endpoints.

Reproduction

See scripts/quantize_nvfp4.py for the full recipe
and QUANTIZATION.md for the end-to-end methodology.

Toolchain: llmcompressor==0.12.0, compressed-tensors==0.17.1,
transformers==5.12.1, torch==2.11.0+cu130, on an NVIDIA GB10 (Blackwell, sm_121).

Related

Notes

  • Needs NVIDIA Blackwell (sm_121, e.g. GB10) for accelerated W4A4; pre-Blackwell GPUs run it weight-only.
  • --reasoning-parser and --tool-call-parser are not auto-detected; pass them explicitly.
  • Thinking mode is on by default; toggle it via the chat template or chat_template_kwargs.

License

Apache-2.0, following the base model. Intended use and all responsibility for use
follow the base model.

Credits

Part of Rogue Quants · NVFP4 component datasheets · collection. Fabricated on GB10 (Blackwell) with llm-compressor. Refusals shown per 100 harmful prompts; "n/a" = not separately measured (base-inherited).

README history 20 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-04uniform schema, no em-dashesdda5a7918.1 KB
    Loading...
  2. 2026-08-04card: uniform characteristics schema (params/active/experts/etc)a37597418.1 KB
    Loading...
  3. 2026-08-03trim redundant Fidelity table (covered by header)9abf3ea16.5 KB
    Loading...
  4. 2026-08-03card: readability - sans labels, mono values, more spacingdd6f14e16.6 KB
    Loading...
  5. 2026-08-03card: fix table borders + readability7535c3f16.6 KB
    Loading...
  6. 2026-08-03card: die sigil + a11y table semantics000eaed16.2 KB
    Loading...
  7. 2026-08-03card: DATASHEET design language58127a515.6 KB
    Loading...
  8. 2026-08-03card: sync post vision/MTP fix8382e5f17.9 KB
    Loading...
  9. 2026-07-07Add Ornith-35B MoE abliterated to Related siblings6ec5baf18.1 KB
    Loading...
  10. 2026-07-04Add Qwopus-27B-v2 abliterated to Related siblings25c918f18 KB
    Loading...
  11. 2026-07-02Add abliterated Qwopus-27B to Related siblings300ad2917.9 KB
    Loading...
  12. 2026-07-01Enrich card: TL;DR, Quickstart (Python+curl), Related, Notes7b7b5ea17.8 KB
    Loading...
  13. 2026-07-01Add Fidelity section372da6d16.1 KB
    Loading...
  14. 2026-07-01Consistency pass: unified banner, remove em dashes, footer1262c5915.7 KB
    Loading...
  15. 2026-07-01Add visual flair to carde39df8c14.6 KB
    Loading...
  16. 2026-06-28nvfp4-40b SWE-bench 14/20 + footprint comparison900649110.2 KB
    Loading...
  17. 2026-06-27Lead with near-lossless headline in the intro67f800910.2 KB
    Loading...
  18. 2026-06-27Restructure card: benchmarks up top, consolidated quality, fresh PPLf13bc8e9.9 KB
    Loading...
  19. 2026-06-26Upload README.md with huggingface_hub3f5056d9.8 KB
    Loading...
  20. 2026-06-24Upload README.md with huggingface_hub37406599.9 KB
    Loading...

Discussions 3 threads

  1. 2026-08-26Thank you for your recipeopen1 💬#3
    Loading...
  2. 2026-08-09https://huggingface.co/DavidAU/Qwen3.6-40B-Fable-Fusion-6-Core-Deckard-Eleanor-…open3 💬#2
    Loading...
  3. 2026-08-01maci0/Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking-NVFP4open5 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration