← back to catalog · registered 2026-09-29 04:57

windowsxp811203/Qwen3.8-27B-Abliterated-NVFP4-v2

windowsxp811203 27B multimodal second-order
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/windowsxp811203%2FQwen3.8-27B-Abliterated-NVFP4-v2"
Response includes
  • classification m-uncensored
  • files 21
  • author_summary 17 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
1
Model age
today
created 2026-09-28

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_5 image-text-to-text nvfp4 fp4 fp8 compressed-tensors vllm abliterated uncensored mtp

Related

Total size
21.4 GB
Files
21
Quantizations
1
Registered
2026-09-29 04:57
Last updated on HF
2026-09-29 04:43

Files by quantization

Auxiliary files 21 files 21.5 GB
model-00001-of-00007.safetensors 3.72 GB c5406e41 download
model-00003-of-00007.safetensors 3.72 GB fdf61793 download
model-00004-of-00007.safetensors 3.72 GB 229f8d9c download
model-00002-of-00007.safetensors 3.72 GB 625df042 download
model-00005-of-00007.safetensors 3.41 GB 9f37ef7e download
model-00006-of-00007.safetensors 2.37 GB 54d83c1d download
model-00007-of-00007.safetensors 810 MB d46385da download
tokenizer.json 12.2 MB 0997f410 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 179 KB d5139fab download
config.json 21.5 KB e327c3bc download
tokenizer_config.json 17.5 KB 5de744b3 download
LICENSE 11.3 KB f938136e download
README.md 10.0 KB e157c791 download
chat_template.jinja 8.74 KB c0c686f9 download
recipe.yaml 1.76 KB 1e91a22b download
SHA256SUMS 1.57 KB 0ce645a8 download
.gitattributes 1.53 KB 52373fe2 download
preprocessor_config.json 390 B 2ea84a43 download
generation_config.json 214 B 0bc3addd download

README current version from Hugging Face


license: apache-2.0
base_model: windowsxp811203/Qwen3.8-27B-Abliterated
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: transformers
tags:

  • nvfp4
  • fp4
  • fp8
  • compressed-tensors
  • vllm
  • qwen3_5
  • abliterated
  • uncensored
  • mtp
  • qwen3.8
  • qwen
  • text-generation
    language:
  • en
  • zh

Qwen3.8-27B-Abliterated — NVFP4 v2 (Gated DeltaNet projections in FP8)

Second quantized build of windowsxp811203/Qwen3.8-27B-Abliterated,
an abliterated (refusal-removed) Qwen/Qwen3.8-27B.

Same NVFP4 treatment of the MLP and full-attention projections as
v1; the difference is that the
three large projections in each of the 48 Gated DeltaNet (linear_attn) layers, which v1 left in bf16
(11.1 GB, 36 % of every decode step's weight traffic), are now FP8-E4M3 per-channel. The small GDN
tensors and everything else are unchanged.

28.6 GB → 23.0 GB on disk, 27.0 → 21.9 GiB loaded¹, ~14 % less time per decode step on the same
card model
(27.1 → 23.5 ms, ≈16 % more tokens/s), paired quality evals inside run-to-run noise, MTP
head intact.

Tested only on Blackwell (sm120: RTX PRO 6000 Blackwell, RTX 5090) with vLLM 0.30.0. vLLM's Marlin
NVFP4 and CUTLASS/Marlin FP8 paths declare sm75+/sm89+ support, but pre-Blackwell cards were not
exercised for this card.

What is and isn't quantized

group v1 v2 (this) count
MLP gate/up/down (64 layers) + full-attention q/k/v/o (16 layers) NVFP4, group 16, float8_e4m3 scales same 256 Linears
linear_attn.in_proj_qkv, in_proj_z, out_proj (48 layers) bf16 FP8-E4M3, per-channel weight scale, dynamic per-token activations (FP8_DYNAMIC) 144 Linears
linear_attn.in_proj_a, in_proj_b, conv1d, norm, A_log, dt_bias bf16 same
mtp.* (draft head) bf16, grafted back after quantization, Linears in ignore same 15 tensors
model.visual.* (vision tower) bf16, bit-identical same 333 keys
lm_head, embeddings bf16 same

config.json carries quantization_config.format = "mixed-precision" with two config_groups
(group_0 = float-quantized for the FP8 layers, group_1 = nvfp4-pack-quantized). vLLM's
compressed-tensors loader routes the NVFP4 group to the Marlin NVFP4 kernel and the FP8 group to
CompressedTensorsW8A8Fp8 (CUTLASS scaled-MM on sm120).

Tensor inventory vs v1: 1,711 → 1,855 tensors (+144 weight_scale); the 144 GDN weights change dtype to
F8_E4M3 with identical shapes; every other tensor has the same name, dtype and shape.

Verification

Paired against v1 — same GPU model (two identical RTX PRO 6000 Blackwell Server cards in one host,
run back-to-back), same engine (vLLM 0.30.0), same flags, same harness
— so for the quality table and
the two Server rows below the only variable is the checkpoint.

Quality (greedy, non-thinking unless noted):

benchmark v1 v2
AdvBench 80-prompt subset, refusals 0/80 0/80
GSM8K, 200 questions, max_tokens 1536 95.5 % 96.5 %
MMLU, 1,000 questions 79.4 % 78.8 %
Needle-in-a-haystack, prompts of 3.1K / 13.6K / 40.7K tokens (plus 1.1K / 4.2K / 6.7K / 20.4K in the functional probe) all retrieved all retrieved
Vision probe (colour of a drawn square), tool call (non-stream + stream), reasoning on/off pass pass

The MMLU gap (−0.6 pt) is one seed's worth of noise at n=1,000. These two columns are comparable only
with each other: not with the v1 card's 77.75 % (400 questions, different prompt/parse) nor with the
parent card's logit-based 82.35 % / 81.10 %. For scale, the parent card's own base-vs-abliterated gap
(−1.05 pp on the full set) is larger than this −0.6 pt.

Speed — single stream, thinking off, 400 output tokens, MTP num_speculative_tokens: 2, fp8 KV
cache, cudagraph_mode: PIECEWISE; per-step time (mean over the three prompts) and decode tok/s and
acceptance from vLLM's own /metrics:

GPU (engine) build ms / step zh prose Python counting draft acceptance
RTX PRO 6000 Blackwell Server (0.30.0, pip wheel) v1 27.1 68 tok/s 106 110 0.43 / 0.95 / 1.00
RTX PRO 6000 Blackwell Server (0.30.0, pip wheel) v2 23.5 77 118 127 0.41 / 0.88 / 1.00
RTX PRO 6000 Blackwell Workstation (0.30.0, Docker; mean of 9 runs) v1 24.5 77 115 122 0.45 / 0.91 / 0.99
RTX 5090 (0.30.0, Docker; same 14001 MHz GDDR7 / 1.79 TB/s as the Workstation card) v2 20.7 89 133 145 0.43 / 0.88 / 1.00

The two Server rows are the paired measurement. The Workstation-v1 and 5090-v2 rows are two different
cards matched only by memory bandwidth (the 5090 run also used a 16K context window, see Usage); they
are shown because that pair is what the Workstation card is expected to do with v2 (memory-bound decode
tracks bytes per step: 31.0 → 25.5 GB with MTP n=2²).

Acceptance is a property of the prompt, not of the build: Chinese free prose sits at ~0.45 on both, code
at ~0.9. The v1 card's headline 76–78 % was measured with num_speculative_tokens: 1 on a different
prompt mix; at n=2 the second draft is accepted less often, so per-token acceptance is lower by
construction. Same build, same prompt, same n gives the same acceptance (see table).

FP8 kernel choice on sm120 makes no measurable difference: CUTLASS W8A8 (default) 23.4–23.6 ms, torch
channel-wise 23.7–23.9 ms, Marlin W8A16 (VLLM_DISABLED_KERNELS=CutlassFP8ScaledMMLinearKernel,ChannelWiseTorchFP8ScaledMMLinearKernel) 23.2–23.3 ms.

¹ Loaded sizes as reported by vLLM 0.30.0's Model loading took; the v1 card's 26.6 GiB was an earlier
engine's figure for the same v1 file.
² Per MTP step with n=2: 64 decoder layers' weights (v1 21.7 GB → v2 16.2 GB) + lm_head 2.54 GB read
three times (target pass + two draft passes) + the MTP block 0.85 GB read twice = 31.0 → 25.5 GB.

Usage

Tested on vLLM 0.30.0 — the pip wheel on the Server cards, the official vllm/vllm-openai:v0.30.0
image (digest sha256:5f5e5352…6d40) on the Workstation/5090 host. Flags used for the RTX PRO 6000 rows:

vllm serve windowsxp811203/Qwen3.8-27B-Abliterated-NVFP4-v2 \
  --max-model-len 262144 --kv-cache-dtype fp8 \
  --enable-prefix-caching --mamba-cache-mode align \
  --speculative-config '{"method":"mtp","num_speculative_tokens":2}' \
  --compilation-config '{"cudagraph_mode":"PIECEWISE"}' \
  --prefix-cache-retention-interval None \
  --attention-config '{"use_trtllm_attention": false}' \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder

Those runs also had --max-num-seqs 224 --gpu-memory-utilization 0.90, which do not affect single-stream
decode. On a 32 GB RTX 5090 the smoke test used --max-model-len 16384 --max-num-seqs 4 --gpu-memory-utilization 0.88 instead: 262K of fp8 KV does not fit next to 21.8 GiB of weights.

Why the --compilation-config, --attention-config and --prefix-cache-retention-interval flags, on
0.30.0 with an SM12x card and fp8 KV (the parser flags are the usual Qwen3 reasoning/tool-call parsers):

  • cudagraph_mode: PIECEWISE + use_trtllm_attention: false — vLLM otherwise auto-selects the
    XQA/TRT-LLM decode kernel with FULL CUDA graphs, a combination reported to silently break long-range
    recall and collapse MTP acceptance (vllm #49010).
  • --prefix-cache-retention-interval None — vLLM 0.30.0 defaults this to 0 for every model; it only
    affects sliding-window/Mamba KV groups — here the 48 GDN layers — where 0 keeps just the
    replay-boundary and shared-prefix checkpoints instead of every block, so multi-turn continuations miss
    the cache on the GDN groups. None restores dense (per-block) retention, the pre-0.30 behaviour.

Thinking is on by default; disable per request with "chat_template_kwargs": {"enable_thinking": false}.
262,144 tokens is the architecture's declared limit; a 262,200-token request is rejected with HTTP 400 as expected.

The compressed-tensors mixed-precision loader is also present in the April-2026 0.19.x nightlies
(source read, not run); only 0.30.0 was run for this card.

Provenance

Quantized with llm-compressor 0.13.0 / compressed-tensors 0.18.0 from the bf16 parent, data-free
(requires_calibration_data: false, as in v1), using two config_groups: NVFP4A16 on
re:.*\.mlp\.(gate|up|down)_proj$ and re:.*\.self_attn\.(q|k|v|o)_proj$, and FP8_DYNAMIC on
re:.*\.linear_attn\.(in_proj_qkv|in_proj_z|out_proj)$; lm_head, embed_tokens, visual, mtp
and the small GDN tensors are in the recipe's ignore. The 15 mtp.* tensors were grafted back in bf16
afterwards and the 8 MTP Linear modules (mtp.fc and the 7 mtp.layers.0 projections) are listed in
quantization_config.ignore — both halves are required; see the v1 card for why a missing ignore entry
reads as 0 % acceptance with clean logs.

The parent was produced by orthogonalizing 131 residual-writing tensors (including embed_tokens)
against a refusal direction at λ=1.5, leaving the vision tower byte-identical. Full recipe and
evaluation in the parent model card.
A llama.cpp build is at
Qwen3.8-27B-Abliterated-GGUF.

SHA256SUMS in this repo lists every file as produced.

Support / 打賞

If these models are useful to you, tips are appreciated — they pay for the GPU time.
如果這些模型對你有幫助,歡迎打賞,用於支應算力成本。

USDT (TRC20) · TPTo32r7vKazpTNaFqfFZ2ztoK1DG88888

Disclaimer

This model will not refuse. It is published for alignment and safety research. You are responsible
for your use of it and for complying with applicable law. Inherits the Apache-2.0 license of the
base model.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-09-29model card: fact-checked revision (paired numbers, needle token counts, flag ...e5d2e0b10 KB
    Loading...
  2. 2026-09-28add README.mdef9aa857.4 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.