← back to catalog · registered 2026-10-05 16:58

zuozijian/Qwen38-Uncensored-NVFP4-Graft

zuozijian multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/zuozijian%2FQwen38-Uncensored-NVFP4-Graft"
Response includes
  • classification m-uncensored
  • files 14
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
1
Model age
today
created 2026-10-05

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_5 image-text-to-text qwen3.8 nvfp4 modelopt tensorrt vllm uncensored abliterated mtp

Related

Total size
17.5 GB
Files
14
Quantizations
1
Registered
2026-10-05 16:58
Last updated on HF
2026-10-05 16:46

Files by quantization

Auxiliary files 14 files 17.5 GB
model.safetensors 16.7 GB f32fc0c9 download
model-mtp-grafted.safetensors 810 MB 90fa0e3e download
tokenizer.json 19.1 MB 06b95093 download
graft-manifest.json 732 KB dc46f74e download
model.safetensors.index.json 197 KB 9f3c6355 download
hf_quant_config.json 54.6 KB 5efe5565 download
config.json 12.7 KB fdd6bca7 download
README.md 9.62 KB 93792c6b download
chat_template.jinja 8.74 KB c0c686f9 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.14 KB 1d134cd2 download
RECOVERY.md 755 B 3288e123 download
generation_config.json 214 B 8b9f95da download

README current version from Hugging Face


license: apache-2.0
base_model: JonathanColetti/Qwen3.8-27B-Uncensored
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: transformers
language:

  • en
  • zh
    tags:
  • qwen3.8
  • nvfp4
  • modelopt
  • tensorrt
  • vllm
  • uncensored
  • abliterated
  • mtp
  • vision

Qwen3.8-27B-Uncensored-NVFP4 (ModelOpt)

NVFP4 quantization of JonathanColetti/Qwen3.8-27B-Uncensored,
produced with NVIDIA TensorRT Model Optimizer 0.43.0
for Blackwell-class inference under vLLM.

~65 GB bf16 → 19.2 GiB. The multi-token-prediction head and the vision tower are
both retained.

What was quantized

400 linear layers to NVFP4 (block size 16, FP8 scales):

Group Modules Quantized
MLP gate_proj / up_proj / down_proj 64 each yes
Full attention q/k/v/o_proj 16 each yes
Gated DeltaNet in_proj_qkv, in_proj_z, out_proj 48 each yes
Gated DeltaNet in_proj_a / in_proj_b 48 each no
Gated DeltaNet conv1d 48 no
Vision tower (model.visual.*) 333 tensors no
lm_head 1 no
MTP head (mtp.*) 15 tensors no

Qwen3.8-27B is a hybrid stack — 64 layers of
3× (Gated DeltaNet → FFN) + 1× (Gated Attention → FFN). The DeltaNet decay and beta
projections (in_proj_a / in_proj_b) are low-rank and precision-sensitive, so they are
left at bf16 along with the causal conv1d.

Fused-layer constraint (important if you re-roll this yourself)

vLLM does not instantiate the DeltaNet input projections separately. It fuses
in_proj_qkv + in_proj_z into a single MergedColumnParallelLinear named
in_proj_qkvz, and in_proj_b + in_proj_a into in_proj_ba. Every shard of a fused
layer must share one precision, or loading aborts during model construction — before a
single weight is read:

ValueError: Detected some but not all shards of
language_model.model.layers.0.linear_attn.in_proj_qkvz are quantized.
All shards of fused layers to have the same precision.

So in_proj_z must be quantized together with in_proj_qkv, even though it is a gate.
Excluding in_proj_a and in_proj_b is fine because they are excluded together, which
leaves in_proj_ba uniform. The same rule applies to qkv_proj and gate_up_proj.

This checkpoint has been verified to satisfy that constraint: every fused group is
internally single-precision, checked per parent module.

exclude_modules naming

exclude_modules is matched against vLLM's module prefixes, by exact string equality
first. Recent transformers emits the checkpoint hierarchy as model.language_model.…,
whereas vLLM builds language_model.model.… — so a config exported verbatim will silently
fail to match, and vLLM will try to quantize layers that have no scales. The exclusion list
here is written in both conventions, and was validated by running vLLM's own
is_layer_skipped / is_layer_excluded over every module in the checkpoint and confirming
its decision matches whether that module actually carries weight_scale tensors.

Calibration

256 samples from garage-bAInd/Open-Platypus,
batch size 16, max sequence length 1024, max calibration. No fine-tuning, no additional
training data.

Serving

vLLM on Blackwell:

vllm serve joshebbs/qwen3.8-27b-uncensored-nvfp4-modelopt \
  --quantization modelopt_fp4 \
  --kv-cache-dtype fp8 \
  --attention-backend flashinfer

The chat template opens a <think> block by default; pass enable_thinking=False to
apply_chat_template for direct answers. Qwen's recommended sampling is
temperature=1.0, top_p=0.95, top_k=20.

Note that transformers cannot load this checkpoint directly — NVFP4 packs two 4-bit
values per byte, so weights are stored at half width and a plain from_pretrained will
report shape mismatches. Use a runtime that understands modelopt_fp4.

Tool calling

The chat template emits tool calls in Qwen's XML dialect
(<tool_call><function=name><parameter=k>v</parameter></function></tool_call>), so vLLM
needs the matching parser. Without both flags, any client sending tool_choice: "auto"
gets 400 "auto" tool choice requires --enable-auto-tool-choice and --tool-call-parser to be set:

--enable-auto-tool-choice --tool-call-parser qwen3_coder

Reasoning parser + max_tokens

With --reasoning-parser qwen3, a response truncated inside the <think> block
(finish_reason: "length") comes back with both content and reasoning_content
empty — the parser needs the closing </think> before it will emit anything. This looks
alarmingly like a corrupted checkpoint but is purely a budget artifact. Either give
thinking mode enough headroom (2500 tokens was still not always enough for a verbose
"explain in detail" prompt) or set enable_thinking=False.

Deployment: 2× NVIDIA DGX Spark (GB10), TP=2

Verified serving on a pair of DGX Spark GB10 nodes joined by a direct 200 Gb/s QSFP link,
tensor-parallel across the two, one GPU per node:

Nodes 2× DGX Spark GB10 (Blackwell, unified memory)
Interconnect direct QSFP, RoCE, 10.10.10.1 ↔ 10.10.10.2
Parallelism -tp 2 --nnodes 2, torch.distributed (no Ray)
vLLM 0.19.2rc1
Weights 19.2 GiB, gpu-memory-utilization 0.75
Context 131072
vllm serve joshebbs/qwen3.8-27b-uncensored-nvfp4-modelopt \
  --served-model-name qwen38-27b \
  --host 0.0.0.0 --port 8000 \
  --max-model-len 131072 \
  --max-num-batched-tokens 8192 \
  --max-num-seqs 4 \
  --gpu-memory-utilization 0.75 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder \
  --reasoning-parser qwen3 \
  --trust-remote-code \
  --dtype auto \
  --kv-cache-dtype fp8 \
  --quantization modelopt_fp4 \
  --attention-backend flashinfer \
  --enable-prefix-caching \
  --enable-chunked-prefill \
  -tp 2 --nnodes 2 --node-rank 0 \
  --master-addr 10.10.10.1 --master-port 29501

Rank 1 runs the same line with --node-rank 1.

Measured startup (warm compile cache)

Phase Time
Distributed init + NCCL ring ~4 min
torch.compile (range 1–8192) 63 s (28 s for the graph)
FlashInfer autotune ~60 s
CUDA graph capture 2 s, 0.58 GiB pool
Engine init total 127 s

First run on a cold FlashInfer cache is far slower — budget 25–30 min. Raise the NCCL
store timeout (300 s default is not enough for multi-node JIT) or rank 1 will drop out
mid-compile while rank 0 is still building kernels.

Measured throughput (TP=2, live server)

vllm bench serve, --dataset-name random --random-input-len 512 --random-output-len 256 --ignore-eos, hitting the running OpenAI chat endpoint. Warm cache, negligible
background load.

Concurrency Output tok/s Total tok/s Mean TTFT Mean ITL
1 20.9 66.9 192 ms 47 ms
8 74.5 238.9 12.4 s 52 ms

Single-user latency is fine (~192 ms first token, ~48 ms per subsequent token). Under
8-way concurrency total-token throughput scales 3.6×; TTFT balloons because prefills
queue against --max-num-seqs 4 / --max-num-batched-tokens 8192, but decode ITL barely
moves — the ceiling is batch admission, not compute. Raise --max-num-seqs if you need
lower TTFT under bursts.

Reproduce:

docker exec -e HF_HUB_OFFLINE=1 vllm_node vllm bench serve \
  --backend openai-chat \
  --base-url http://localhost:8000 --endpoint /v1/chat/completions \
  --model qwen38-27b \
  --tokenizer /root/.cache/huggingface/hub/models--joshebbs--qwen3.8-27b-uncensored-nvfp4-modelopt/snapshots/<hash> \
  --dataset-name random --random-input-len 512 --random-output-len 256 \
  --num-prompts 32 --max-concurrency 8 --ignore-eos

Notes specific to this stack

  • Both nodes must hold the same checkpoint revision. vLLM resolves the repo id to a
    local snapshot path per node; if one node's HF cache is a revision behind, each rank
    silently loads different weights and the run hangs in distributed init rather than
    reporting a mismatch. Check refs/main on both.
  • Prefix caching puts the Mamba/DeltaNet cache in align mode, which vLLM flags as
    experimental for this architecture. Drop --enable-prefix-caching first if you see
    output corruption.
  • The MTP head ships in the checkpoint but is not loaded unless you configure speculative
    decoding; speculative_config=None leaves model-mtp-grafted.safetensors unused.

Refusal behaviour — inherited, not re-measured

The base checkpoint's author reports 12/100 refusals vs 98/100 for stock Qwen3.8-27B
on the test split of mlabonne/harmful_behaviors,
measured in non-thinking mode, using Heretic
(200-trial search co-minimizing refusal count against KL divergence from base).

Those numbers describe the bf16 source, not this quantization. The refusal edit lives
in o_proj and down_proj, which are exactly the tensors compressed here, so the effect
could in principle be attenuated. The bf16 source was spot-checked as compliant before
quantization.

Refusals are reduced, not eliminated, in the source model.

Credits

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration