license: apache-2.0
base_model: windowsxp811203/Qwen3.8-27B-Abliterated
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: transformers
tags:
- nvfp4
- fp4
- fp8
- compressed-tensors
- vllm
- qwen3_5
- abliterated
- uncensored
- mtp
- qwen3.8
- qwen
- text-generation
language: - en
- zh
Qwen3.8-27B-Abliterated — NVFP4 v2 (Gated DeltaNet projections in FP8)
Second quantized build of windowsxp811203/Qwen3.8-27B-Abliterated,
an abliterated (refusal-removed) Qwen/Qwen3.8-27B.
Same NVFP4 treatment of the MLP and full-attention projections as
v1; the difference is that the
three large projections in each of the 48 Gated DeltaNet (linear_attn) layers, which v1 left in bf16
(11.1 GB, 36 % of every decode step's weight traffic), are now FP8-E4M3 per-channel. The small GDN
tensors and everything else are unchanged.
28.6 GB → 23.0 GB on disk, 27.0 → 21.9 GiB loaded¹, ~14 % less time per decode step on the same
card model (27.1 → 23.5 ms, ≈16 % more tokens/s), paired quality evals inside run-to-run noise, MTP
head intact.
Tested only on Blackwell (sm120: RTX PRO 6000 Blackwell, RTX 5090) with vLLM 0.30.0. vLLM's Marlin
NVFP4 and CUTLASS/Marlin FP8 paths declare sm75+/sm89+ support, but pre-Blackwell cards were not
exercised for this card.
What is and isn't quantized
| group | v1 | v2 (this) | count |
|---|---|---|---|
MLP gate/up/down (64 layers) + full-attention q/k/v/o (16 layers) |
NVFP4, group 16, float8_e4m3 scales |
same | 256 Linears |
linear_attn.in_proj_qkv, in_proj_z, out_proj (48 layers) |
bf16 | FP8-E4M3, per-channel weight scale, dynamic per-token activations (FP8_DYNAMIC) |
144 Linears |
linear_attn.in_proj_a, in_proj_b, conv1d, norm, A_log, dt_bias |
bf16 | same | |
mtp.* (draft head) |
bf16, grafted back after quantization, Linears in ignore |
same | 15 tensors |
model.visual.* (vision tower) |
bf16, bit-identical | same | 333 keys |
lm_head, embeddings |
bf16 | same |
config.json carries quantization_config.format = "mixed-precision" with two config_groups
(group_0 = float-quantized for the FP8 layers, group_1 = nvfp4-pack-quantized). vLLM's
compressed-tensors loader routes the NVFP4 group to the Marlin NVFP4 kernel and the FP8 group toCompressedTensorsW8A8Fp8 (CUTLASS scaled-MM on sm120).
Tensor inventory vs v1: 1,711 → 1,855 tensors (+144 weight_scale); the 144 GDN weights change dtype toF8_E4M3 with identical shapes; every other tensor has the same name, dtype and shape.
Verification
Paired against v1 — same GPU model (two identical RTX PRO 6000 Blackwell Server cards in one host,
run back-to-back), same engine (vLLM 0.30.0), same flags, same harness — so for the quality table and
the two Server rows below the only variable is the checkpoint.
Quality (greedy, non-thinking unless noted):
| benchmark | v1 | v2 |
|---|---|---|
| AdvBench 80-prompt subset, refusals | 0/80 | 0/80 |
GSM8K, 200 questions, max_tokens 1536 |
95.5 % | 96.5 % |
| MMLU, 1,000 questions | 79.4 % | 78.8 % |
| Needle-in-a-haystack, prompts of 3.1K / 13.6K / 40.7K tokens (plus 1.1K / 4.2K / 6.7K / 20.4K in the functional probe) | all retrieved | all retrieved |
| Vision probe (colour of a drawn square), tool call (non-stream + stream), reasoning on/off | pass | pass |
The MMLU gap (−0.6 pt) is one seed's worth of noise at n=1,000. These two columns are comparable only
with each other: not with the v1 card's 77.75 % (400 questions, different prompt/parse) nor with the
parent card's logit-based 82.35 % / 81.10 %. For scale, the parent card's own base-vs-abliterated gap
(−1.05 pp on the full set) is larger than this −0.6 pt.
Speed — single stream, thinking off, 400 output tokens, MTP num_speculative_tokens: 2, fp8 KV
cache, cudagraph_mode: PIECEWISE; per-step time (mean over the three prompts) and decode tok/s and
acceptance from vLLM's own /metrics:
| GPU (engine) | build | ms / step | zh prose | Python | counting | draft acceptance |
|---|---|---|---|---|---|---|
| RTX PRO 6000 Blackwell Server (0.30.0, pip wheel) | v1 | 27.1 | 68 tok/s | 106 | 110 | 0.43 / 0.95 / 1.00 |
| RTX PRO 6000 Blackwell Server (0.30.0, pip wheel) | v2 | 23.5 | 77 | 118 | 127 | 0.41 / 0.88 / 1.00 |
| RTX PRO 6000 Blackwell Workstation (0.30.0, Docker; mean of 9 runs) | v1 | 24.5 | 77 | 115 | 122 | 0.45 / 0.91 / 0.99 |
| RTX 5090 (0.30.0, Docker; same 14001 MHz GDDR7 / 1.79 TB/s as the Workstation card) | v2 | 20.7 | 89 | 133 | 145 | 0.43 / 0.88 / 1.00 |
The two Server rows are the paired measurement. The Workstation-v1 and 5090-v2 rows are two different
cards matched only by memory bandwidth (the 5090 run also used a 16K context window, see Usage); they
are shown because that pair is what the Workstation card is expected to do with v2 (memory-bound decode
tracks bytes per step: 31.0 → 25.5 GB with MTP n=2²).
Acceptance is a property of the prompt, not of the build: Chinese free prose sits at ~0.45 on both, code
at ~0.9. The v1 card's headline 76–78 % was measured with num_speculative_tokens: 1 on a different
prompt mix; at n=2 the second draft is accepted less often, so per-token acceptance is lower by
construction. Same build, same prompt, same n gives the same acceptance (see table).
FP8 kernel choice on sm120 makes no measurable difference: CUTLASS W8A8 (default) 23.4–23.6 ms, torch
channel-wise 23.7–23.9 ms, Marlin W8A16 (VLLM_DISABLED_KERNELS=CutlassFP8ScaledMMLinearKernel,ChannelWiseTorchFP8ScaledMMLinearKernel) 23.2–23.3 ms.
¹ Loaded sizes as reported by vLLM 0.30.0's Model loading took; the v1 card's 26.6 GiB was an earlier
engine's figure for the same v1 file.
² Per MTP step with n=2: 64 decoder layers' weights (v1 21.7 GB → v2 16.2 GB) + lm_head 2.54 GB read
three times (target pass + two draft passes) + the MTP block 0.85 GB read twice = 31.0 → 25.5 GB.
Usage
Tested on vLLM 0.30.0 — the pip wheel on the Server cards, the official vllm/vllm-openai:v0.30.0
image (digest sha256:5f5e5352…6d40) on the Workstation/5090 host. Flags used for the RTX PRO 6000 rows:
vllm serve windowsxp811203/Qwen3.8-27B-Abliterated-NVFP4-v2 \
--max-model-len 262144 --kv-cache-dtype fp8 \
--enable-prefix-caching --mamba-cache-mode align \
--speculative-config '{"method":"mtp","num_speculative_tokens":2}' \
--compilation-config '{"cudagraph_mode":"PIECEWISE"}' \
--prefix-cache-retention-interval None \
--attention-config '{"use_trtllm_attention": false}' \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder
Those runs also had --max-num-seqs 224 --gpu-memory-utilization 0.90, which do not affect single-stream
decode. On a 32 GB RTX 5090 the smoke test used --max-model-len 16384 --max-num-seqs 4 --gpu-memory-utilization 0.88 instead: 262K of fp8 KV does not fit next to 21.8 GiB of weights.
Why the --compilation-config, --attention-config and --prefix-cache-retention-interval flags, on
0.30.0 with an SM12x card and fp8 KV (the parser flags are the usual Qwen3 reasoning/tool-call parsers):
cudagraph_mode: PIECEWISE+use_trtllm_attention: false— vLLM otherwise auto-selects the
XQA/TRT-LLM decode kernel with FULL CUDA graphs, a combination reported to silently break long-range
recall and collapse MTP acceptance (vllm #49010).--prefix-cache-retention-interval None— vLLM 0.30.0 defaults this to0for every model; it only
affects sliding-window/Mamba KV groups — here the 48 GDN layers — where0keeps just the
replay-boundary and shared-prefix checkpoints instead of every block, so multi-turn continuations miss
the cache on the GDN groups.Nonerestores dense (per-block) retention, the pre-0.30 behaviour.
Thinking is on by default; disable per request with "chat_template_kwargs": {"enable_thinking": false}.
262,144 tokens is the architecture's declared limit; a 262,200-token request is rejected with HTTP 400 as expected.
The compressed-tensors mixed-precision loader is also present in the April-2026 0.19.x nightlies
(source read, not run); only 0.30.0 was run for this card.
Provenance
Quantized with llm-compressor 0.13.0 / compressed-tensors 0.18.0 from the bf16 parent, data-free
(requires_calibration_data: false, as in v1), using two config_groups: NVFP4A16 onre:.*\.mlp\.(gate|up|down)_proj$ and re:.*\.self_attn\.(q|k|v|o)_proj$, and FP8_DYNAMIC onre:.*\.linear_attn\.(in_proj_qkv|in_proj_z|out_proj)$; lm_head, embed_tokens, visual, mtp
and the small GDN tensors are in the recipe's ignore. The 15 mtp.* tensors were grafted back in bf16
afterwards and the 8 MTP Linear modules (mtp.fc and the 7 mtp.layers.0 projections) are listed inquantization_config.ignore — both halves are required; see the v1 card for why a missing ignore entry
reads as 0 % acceptance with clean logs.
The parent was produced by orthogonalizing 131 residual-writing tensors (including embed_tokens)
against a refusal direction at λ=1.5, leaving the vision tower byte-identical. Full recipe and
evaluation in the parent model card.
A llama.cpp build is at
Qwen3.8-27B-Abliterated-GGUF.
SHA256SUMS in this repo lists every file as produced.
Support / 打賞
If these models are useful to you, tips are appreciated — they pay for the GPU time.
如果這些模型對你有幫助,歡迎打賞,用於支應算力成本。
USDT (TRC20) · TPTo32r7vKazpTNaFqfFZ2ztoK1DG88888
Disclaimer
This model will not refuse. It is published for alignment and safety research. You are responsible
for your use of it and for complying with applicable law. Inherits the Apache-2.0 license of the
base model.