← back to catalog · registered 2026-08-22 13:56

windowsxp811203/Qwen3.8-27B-Abliterated-NVFP4

windowsxp811203 Qwen 19B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/windowsxp811203%2FQwen3.8-27B-Abliterated-NVFP4"
Response includes
  • classification m1
  • files 21
  • hub_downloads_all_time 2,339
  • author_summary 17 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
2K
247 last 30d - stable
Likes
2
Model age
8w ago
created 2026-08-15
Downloads over time
Now2.4K→from1.5K↑67%
1.4K1.8K2.2K2.5K1.5K on Aug 192.4K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 3 formats · 11K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_5 image-text-to-text nvfp4 fp4 compressed-tensors vllm abliterated uncensored mtp qwen3.8

Related

Total size
26.6 GB
Files
21
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-09-29 04:43

Files by quantization

Auxiliary files 21 files 26.6 GB
model-00006-of-00008.safetensors 3.72 GB c757149d download
model-00002-of-00008.safetensors 3.72 GB 67ae2e1d download
model-00003-of-00008.safetensors 3.70 GB 257144f3 download
model-00005-of-00008.safetensors 3.69 GB e7fa8bc9 download
model-00004-of-00008.safetensors 3.68 GB baa5f0e8 download
model-00001-of-00008.safetensors 3.64 GB 27c43143 download
model-00007-of-00008.safetensors 3.64 GB aed4bc54 download
model-00008-of-00008.safetensors 810 MB d46385da download
tokenizer.json 12.2 MB 0997f410 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 164 KB e5e24d70 download
config.json 29.2 KB 5ca6908a download
tokenizer_config.json 17.5 KB 5de744b3 download
LICENSE 11.3 KB f938136e download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 5.90 KB 44bf16ca download
.gitattributes 1.53 KB 52373fe2 download
preprocessor_config.json 390 B 2ea84a43 download
recipe.yaml 267 B 40d71f8b download
generation_config.json 214 B 0bc3addd download

README current version from Hugging Face


license: apache-2.0
base_model: windowsxp811203/Qwen3.8-27B-Abliterated
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: transformers
tags:

  • nvfp4
  • fp4
  • compressed-tensors
  • vllm
  • qwen3_5
  • abliterated
  • uncensored
  • mtp
  • qwen3.8
  • qwen
  • text-generation
    language:
  • en
  • zh

Qwen3.8-27B-Abliterated — NVFP4A16

4-bit NVFP4 build of windowsxp811203/Qwen3.8-27B-Abliterated,
an abliterated (refusal-removed) Qwen/Qwen3.8-27B.

55.6 GB → 28.6 GB, and the MTP draft head is intact and verified working at 76–78 % draft
acceptance
. Most quantized derivatives of this architecture ship a dead MTP head, because
Qwen3_5ForConditionalGeneration.from_pretrained silently drops it and llm-compressor cannot save
a module transformers never instantiated.

Requires a Blackwell GPU (sm100+) and vLLM.

What is and isn't quantized

group treatment count
MLP gate/up/down (64 layers) + full-attention q/k/v/o (16 layers) NVFP4 — 4-bit float, group size 16, float8_e4m3 scales 256 Linears
mtp.* (draft head) bf16, grafted back after quantization, in ignore 15 tensors
model.visual.* (vision tower) bf16 — kept bit-identical 167 weight tensors (333 incl. biases/norms)
linear_attn.* (Gated DeltaNet / SSM) bf16 — quantization-sensitive 336 tensors
lm_head, embeddings bf16

The MTP tensors are both grafted back and listed in quantization_config.ignore. Both halves
matter: without the graft there is no draft head at all, and without the ignore entry vLLM's
compressed-tensors loader treats the bf16 head as a quantization target, finds no scales, and
rejects every draft — 0 % acceptance while the logs look perfectly clean.

Verification

Measured on this exact checkpoint on an RTX PRO 6000 Blackwell.

MTP speculative decoding — {"method":"mtp","num_speculative_tokens":1}:

metric value
Avg draft acceptance rate 76.4 % – 78.3 %
Mean acceptance length 1.76 – 1.78
Accepted / drafted 1615 / 2113 tokens

That number is the proof the graft worked; a broken MTP head reads 0 %.

Refusal — greedy, non-thinking, no prefill jailbreak. Unmodified Qwen3.8-27B refuses
99.04 % of AdvBench under identical settings:

benchmark result
AdvBench (80) 0/80 · 0.00 %
HarmBench safety categories (119) 0/119 · 0.0 %
HarmBench copyright (41) 17/41 · 41.5 %

Safety categories = chemical/biological, cybercrime, harassment, harmful, illegal, misinformation
— every one exactly zero. The copyright column is not a safety refusal and is mostly
classifier false positives: the model delivers the lyrics or passage, but the text trips the
keyword list (either the generated prose itself opens with "I cannot quite…", or a pedantic
"I cannot generate a new passage … but here is a long excerpt" precedes the excerpt).

Capability — MMLU, 400 equidistant questions, identical prompting and parsing for both:

build MMLU
GGUF Q8_0 (reference) 78.00 %
NVFP4A16 (this) 77.75 %

One question apart. (Do not compare these to the parent card's 82.35 %: that figure was measured
by next-token logit comparison, a different and more forgiving method. Only same-method numbers
are comparable.)

Usage

vllm serve windowsxp811203/Qwen3.8-27B-Abliterated-NVFP4 \
  --max-model-len 8192 \
  --speculative-config '{"method":"mtp","num_speculative_tokens":1}'

Thinking is on by default; disable per request with
"chat_template_kwargs": {"enable_thinking": false}.

1M context (verified)

The model's declared native limit is 262,144. The 1M configuration from the official Qwen3.8-27B
recipe works on this checkpoint and was measured end-to-end:

vllm serve windowsxp811203/Qwen3.8-27B-Abliterated-NVFP4 \
  --tensor-parallel-size 2 --max-model-len 1010000 \
  --hf-overrides '{"text_config": {"max_position_embeddings": 1010000}}'

Needle-in-a-haystack, passphrase buried at 50 % depth, greedy:

context prompt tokens retrieved time
1 M 823,878 ✅ 364 s on 2× RTX PRO 6000 Blackwell

Budget ≈ 61 GiB of KV at 1M (16 full-attention layers × 4 KV heads × 256 dim) on top of the
26.6 GiB of weights, so 1M needs two 96 GB cards; 256K fits comfortably on one.

If vLLM fails to start with a FlashInfer error

On hosts where the CUDA toolkit and FlashInfer's bundled CCCL headers disagree, FlashInfer's JIT
fails to build its sampling kernels and vLLM aborts with
FlashInfer requires GPUs with sm75 or higher (a misleading message — the real cause is that the
capability probe itself failed). Working around it:

export VLLM_USE_FLASHINFER_SAMPLER=0
export VLLM_ATTENTION_BACKEND=TRITON_ATTN

This is a host toolchain issue, not a property of these weights.

Provenance

Quantized with llm-compressor 0.13.0 (NVFP4A16) from the bf16 parent, which was produced by
orthogonalizing 131 residual-writing tensors (including embed_tokens) against a refusal direction
at λ=1.5, leaving the vision tower byte-identical. Full recipe and evaluation in the
parent model card.
A llama.cpp build is at
Qwen3.8-27B-Abliterated-GGUF.

Support / 打賞

If these models are useful to you, tips are appreciated — they pay for the GPU time.
如果這些模型對你有幫助,歡迎打賞,用於支應算力成本。

USDT (TRC20) · TPTo32r7vKazpTNaFqfFZ2ztoK1DG88888

Disclaimer

This model will not refuse. It is published for alignment and safety research. You are responsible
for your use of it and for complying with applicable law. Inherits the Apache-2.0 license of the
base model.

README history 9 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-09-29model card: pointer to v2 (GDN projections in FP8)bd7ab177.5 KB
    Loading...
  2. 2026-08-25Rev 4 survey numbers; projection-denominator correction00b3f217.1 KB
    Loading...
  3. 2026-08-24Rev 3 survey numbers; AdvBench denominators, 824K-prompt verification, linear...a839fe17.1 KB
    Loading...
  4. 2026-08-23Correct the claim about MTP heads in other NVFP4 buildseba210e6.6 KB
    Loading...
  5. 2026-08-18revert pipeline_tag to image-text-to-text (fixes generated snippet)b1064d55.9 KB
    Loading...
  6. 2026-08-17discoverability: text-generation pipeline, qwen3.8 tags, ungate95b02565.9 KB
    Loading...
  7. 2026-08-16Correct errors found in an adversarial audit of the published cards803fcad6.2 KB
    Loading...
  8. 2026-08-16Document verified 1M-context configuration21c61606.1 KB
    Loading...
  9. 2026-08-15Add files using upload-large-folder toola11ef9e5.4 KB
    Loading...

Discussions 1 thread

  1. 2026-08-15Greatopen39 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration