← back to catalog · registered 2026-08-23 21:02

AudioJaguar/Qwen3.8-27B-Uncensored-NVFP4-W4A4

AudioJaguar Qwen 27B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/AudioJaguar%2FQwen3.8-27B-Uncensored-NVFP4-W4A4"
Response includes
  • classification m-uncensored
  • files 22
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
159
Likes
1
Model age
6w ago
created 2026-08-23

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text nvfp4 modelopt sglang dflash2 speculative-decoding dgx-spark gb10 text-generation

Related

Total size
20.4 GB
Files
22
Quantizations
1
Registered
2026-08-23 21:02

Files by quantization

Auxiliary files 22 files 20.4 GB
model-00002-of-00003.safetensors 9.30 GB 66acbc87 download
model-00001-of-00003.safetensors 9.28 GB 2f9578a6 download
model-00003-of-00003.safetensors 1.83 GB 3b799cd0 download
tokenizer.json 12.2 MB 0997f410 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
.quant_summary.txt 307 KB c63cdc56 download
quant_summary.txt 307 KB c63cdc56 download
model.safetensors.index.json 210 KB 146d8eb2 download
config.json 85.6 KB 33f7e788 download
hf_quant_config.json 52.5 KB 36f423d5 download
LICENSE 11.3 KB f938136e download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 4.05 KB 3e81ce20 download
orca-nvfp4-w4a4-fp8attn.yaml 1.69 KB fe831506 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.09 KB 77053000 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 214 B 0bc3addd download
modelopt-commit.txt 52.0 B a2f343b4 download

README current version from Hugging Face


license: apache-2.0
base_model: orcarouter/Qwen3.8-27B-Uncensored
base_model_relation: quantized
pipeline_tag: text-generation
library_name: transformers
tags:

  • nvfp4
  • modelopt
  • sglang
  • dflash2
  • speculative-decoding
  • dgx-spark
  • gb10

Qwen3.8-27B-Uncensored-NVFP4-W4A4

NVFP4 quantization of orcarouter/Qwen3.8-27B-Uncensored,
built to be served with DFlash2 speculative decoding on a single bandwidth-limited
Blackwell GPU. The layout matches RadixArk/Qwen3.8-27B-NVFP4 (revision 52d1adc).
21 GB on disk.

Why this layout

DFlash2 drafts a block of tokens per step and the target verifies them. That makes decode
even more bandwidth-bound than usual, and it adds one specific cost: the drafter's
candidate selector reads the target's lm_head a second time every step, on top of
the verify pass. A BF16 head is 2.5 GB; NVFP4 is 0.7 GB. Paid twice per step, that
difference is worth real throughput on a memory-bandwidth-limited card.

So everything that can be NVFP4 is NVFP4 — all 64 layers' MLP projections and the output
head — with FP8 reserved for the attention and linear-attention projections.

Layout

Modules Format
All 64 layers' MLP gate_proj / up_proj / down_proj, plus lm_head NVFP4 W4A4, group 16 (193 modules)
Self-attention q/k/v/o_proj; Gated DeltaNet in_proj_qkv / in_proj_z / out_proj FP8 W8A8 (208 modules)
KV cache FP8 scales
Vision tower, MTP head, conv1d, in_proj_a/b, embeddings, norms BF16

hf_quant_config.json's per-layer map is byte-identical to RadixArk's pinned revision.
The MTP head (15 tensors) is preserved unquantized and listed in exclude_modules, so
MTP speculative decoding remains available as a fallback if you'd rather not run DFlash2.

Serving with DFlash2

Pairs with the z-lab/Qwen3.8-27B-DFlash2
drafter (BF16, ~3.85 GB). SGLang flags:

--speculative-algorithm DFLASH \
--speculative-draft-model-path z-lab/Qwen3.8-27B-DFlash2 \
--speculative-num-draft-tokens 8 \
--speculative-draft-model-quantization unquant

Keep the drafter unquantized — a quantized drafter loses more in acceptance than it saves
in bandwidth, and quantized-drafter support is still incomplete upstream.

DFlash2 support is recent. At time of writing it requires a build carrying SGLang PR
#35371 plus a fix allowing the candidate selector to run through a quantized lm_head
— which this checkpoint has. The failure mode when an engine can't is usually a crash or
silently disabled speculation, not wrong output. Speculative decoding is lossless: the
target verifies every drafted token, so acceptance affects speed only, never output.

The lm_head trade-off

A quantized lm_head shifts the candidate selector's top-k slightly, which can lower
draft acceptance. RadixArk moved their own checkpoint's head back to BF16 for that reason
— a sensible call on a GB300, where 3.4 GB/step is noise. On a GB10 that same 3.4 GB is
about 15% of the step, and the NVFP4-head layout came out faster end-to-end despite any
acceptance cost. Which side you land on depends on your memory bandwidth. If you want a
BF16-head variant, the recipe in this repo has a comment showing the two-line change.

Reproduction

  • NVIDIA Model Optimizer, commit in modelopt-commit.txt
  • Recipe: orca-nvfp4-w4a4-fp8attn.yaml (in this repo)
  • Calibration: cnn_dailymail 3.0.0, 1024 samples, seq len 512, batch 8, max-calibration
  • Per-module result: quant_summary.txt

Caveats

  • Requires Blackwell (NVFP4 tensor cores). Built and tested on GB10 (sm_121).
  • This is a quantization of an abliterated model: the base has had refusal behaviour
    removed and will not decline requests the original Qwen3.8-27B would. Safety
    characteristics come from the base model, not from this quantization. Use accordingly.

License and attribution

Apache 2.0, inherited from Qwen/Qwen3.8-27B via orcarouter/Qwen3.8-27B-Uncensored.
Layout follows RadixArk/Qwen3.8-27B-NVFP4. Drafter by z-lab.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration