← back to catalog · registered 2026-08-22 13:56

zebulon-prime/Qwen3.8-27B-Dominatrix-abliterated-MTP-NVFP4

zebulon-prime Qwen 8.6B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/zebulon-prime%2FQwen3.8-27B-Dominatrix-abliterated-MTP-NVFP4"
Response includes
  • classification m1
  • files 14
  • hub_downloads_all_time 118
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
118
Likes
0
Model age
7w ago
created 2026-08-21
Downloads over time
Now153→from32↑378%
267211916532 on Aug 19153 on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
transformers safetensors qwen3_5 image-text-to-text nvfp4 fp4 modelopt tensorrt sglang dflash2 speculative-decoding roleplay

Related

Total size
22.1 GB
Files
14
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-21 00:43

Files by quantization

Auxiliary files 14 files 22.1 GB
model-00002-of-00003.safetensors 9.30 GB 36ba0539 download
model-00001-of-00003.safetensors 9.28 GB 2f2aa0ad download
model-00003-of-00003.safetensors 3.54 GB 4d8c788c download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 210 KB 5e98215c download
config.json 85.4 KB 9c205dae download
hf_quant_config.json 52.4 KB 7cc93bba download
LICENSE 11.1 KB d6456956 download
chat_template.jinja 8.97 KB 675afe7c download
README.md 3.53 KB 548fafd7 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.13 KB f2d9f03f download
generation_config.json 214 B 3f25ead4 download

README current version from Hugging Face


license: apache-2.0
language:

  • en
    library_name: transformers
    pipeline_tag: image-text-to-text
    base_model:
  • zebulon-prime/Qwen3.8-27B-Dominatrix-abliterated
    base_model_relation: quantized
    tags:
  • nvfp4
  • fp4
  • modelopt
  • tensorrt
  • sglang
  • dflash2
  • speculative-decoding
  • roleplay
  • creative-writing
  • abliterated
  • uncensored
  • qwen3
  • not-for-all-audiences

Qwen3.8-27B-Dominatrix-abliterated — MTP-NVFP4 (DFlash 2 compatible)

Mixed-precision NVFP4/FP8 PTQ of
zebulon-prime/Qwen3.8-27B-Dominatrix-abliterated
— allura-org's Dominatrix roleplay finetune with huihui-ai's refusal direction projected out.

23 GB, built with NVIDIA TensorRT Model Optimizer for Blackwell (SM120) inference in
SGLang or vLLM.

The distinguishing feature: lm_head is left dense BF16, which is a hard requirement for
DFlash 2 speculative decoding. Most ModelOpt NVFP4 exports of this architecture quantize
lm_head and therefore cannot run DFlash 2 at all. Cost of the dense head is ~1.9 GB of
VRAM over a packed one.


Quantization layout

component precision
MLP gate_proj / up_proj / down_proj NVFP4 W4A4
self_attn q/k/v/o, linear_attn projections FP8 e4m3
KV cache FP8
lm_head BF16, dense
embed_tokens, MTP head, vision tower BF16

Export format is ModelOpt MIXED_PRECISION with a per-layer map in hf_quant_config.json.
Calibrated on in-domain ChatML roleplay text rather than a generic news corpus.

hf_quant_config.json records producer.version: 0.0.0 because it was built from an
editable install. That field is not meaningful provenance.


Serving

SGLang with DFlash 2

Requires the z-lab/Qwen3.8-27B-DFlash2
drafter and an SGLang build including PR #35371.

sglang serve \
  --trust-remote-code \
  --model-path /models/Qwen3.8-27B-Dominatrix-abliterated-MTP-NVFP4 \
  --mem-fraction-static 0.70 \
  --attention-backend flashinfer \
  --chunked-prefill-size 2048 \
  --reasoning-parser qwen3 \
  --speculative-algorithm DFLASH \
  --speculative-draft-model-path /models/Qwen3.8-27B-DFlash2-zlab \
  --speculative-dflash-block-size 8 \
  --speculative-draft-model-quantization unquant

--kv-cache-dtype can be omitted: this checkpoint declares kv_cache_quant_algo: FP8, so
SGLang's auto resolves it.

Without speculative decoding

Drop the four --speculative-* flags. The dense lm_head buys nothing in that configuration
but is otherwise harmless.

MTP

The mtp.* tensors survive the quant in BF16, so MTP speculation remains available as an
alternative drafter. Pick one — MTP or DFlash 2, not both.

Hardware

NVFP4 requires Blackwell (SM120+) for native FP4 tensor-core execution. Built and tested
on an RTX PRO 6000 Blackwell. Weights are ~23 GB, leaving room for a long-context KV cache and
the 2B DFlash 2 drafter.

Quality

Fidelity of the underlying BF16 abliteration versus stock Dominatrix is summarised on the
BF16 card.

The quantization error of this build has not been measured. Treat it as unquantified.
Leaving lm_head dense should help, since the output projection is among the most
quantization-sensitive layers, but that is reasoning, not a measurement.

Sampler guidance from upstream Dominatrix carries over: temperature 1.0–1.25 with min_p 0.1 or top_p 0.95; some prefer 0.7 and nothing else.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-21Initial commit24390a03.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration