← back to catalog · registered 2026-08-22 13:56

preetpatel/Qwen3.8-27B-Uncensored-NVFP4

preetpatel Qwen 24B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/preetpatel%2FQwen3.8-27B-Uncensored-NVFP4"
Response includes
  • classification m1
  • files 12
  • hub_downloads_all_time 4,747
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
5K
1K last 30d - stable
Likes
2
Model age
7w ago
created 2026-08-19
Downloads over time
Now5.2K→from229↑2,185%
01.9K3.8K5.7K229 on Aug 195.2K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
safetensors qwen3_5 qwen3.8 nvfp4 fp4 compressed-tensors vllm blackwell rtx-5090 uncensored abliterated quantized

Related

Total size
18.4 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-19 06:45

Files by quantization

Auxiliary files 12 files 18.4 GB
model.safetensors 18.4 GB bc642c11 download
tokenizer.json 19.1 MB 06b95093 download
config.json 14.7 KB 0189670c download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 4.80 KB 800f83af download
.gitattributes 1.71 KB d0b7963b download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.14 KB b4acebe0 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
recipe.yaml 239 B 67204882 download
generation_config.json 214 B 0bc3addd download

README current version from Hugging Face


license: apache-2.0
base_model: orcarouter/Qwen3.8-27B-Uncensored
base_model_relation: quantized
language:

  • en
  • zh
    pipeline_tag: image-text-to-text
    tags:
  • qwen3.8
  • qwen3_5
  • nvfp4
  • fp4
  • compressed-tensors
  • vllm
  • blackwell
  • rtx-5090
  • uncensored
  • abliterated
  • quantized
  • vision-language
  • function-calling
  • reasoning
  • coding-agent

Qwen3.8-27B-Uncensored-NVFP4

NVFP4 (W4A4) quantization of
orcarouter/Qwen3.8-27B-Uncensored,
an abliterated fine-tune of Qwen/Qwen3.8-27B.
~19GB (down from ~56GB BF16) — built for single consumer Blackwell GPUs.
On an RTX 5090 (32GB) it serves with 160k context and room to spare.

  • FP4 weights and FP4 activations on the transformer linear layers →
    native speedups on Blackwell (SM120: RTX 50-series, RTX PRO; SM100: B200)
  • Kept in BF16: lm_head, the vision tower, and all gated-DeltaNet
    linear_attn layers (the hybrid architecture's linear-attention blocks)
  • Vision, tool calling, and reasoning (<think>) all functional

Serving with vLLM

Verified config for a single RTX 5090 (32GB), tested on vLLM 0.27.1:

vllm serve preetpatel/Qwen3.8-27B-Uncensored-NVFP4 \
    --max-model-len 163840 \
    --kv-cache-dtype fp8 \
    --gpu-memory-utilization 0.87 \
    --max-num-seqs 8 \
    --max-num-batched-tokens 8192 \
    --limit-mm-per-prompt '{"image": 2, "video": 0}' \
    --mm-processor-kwargs '{"max_pixels": 2097152}' \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder \
    --reasoning-parser qwen3

Notes from real-world testing on the 5090:

  • Leave headroom. A tighter config (BF16 KV, 131k context,
    higher utilization) boots fine but OOMs mid-request: the gated-DeltaNet
    prefill kernel makes ~200MB transient allocations that the startup
    profiler does not fully account for. --gpu-memory-utilization 0.87
    leaves ~1.5GB of slack; PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
    helps against fragmentation.
  • FP8 KV cache halves KV to ~32KiB/token. The hybrid architecture only
    keeps KV for 16 of 64 layers, so 160k context needs only ~5GB of KV.
  • Tool calls use the Qwen3-Coder XML format
    (<function=...><parameter=...>) — --tool-call-parser qwen3_coder
    is required for structured tool_calls output; --reasoning-parser qwen3
    moves <think> content into reasoning_content.
  • The model's generation_config.json supplies the recommended sampling
    defaults (temperature 1.0, top_p 0.95, top_k 20).

Measured results (RTX 5090, vLLM 0.27.1)

Test Result
Needle retrieval, 155k-token prompt, needle at 85% depth exact
Needle retrieval, 155k-token prompt, needle at 10% depth exact
2 × 60k-token prompts, concurrent both exact, 16s wall
Structured tool calling (tools API) correct tool_calls + finish_reason
Prefill throughput ~6-7k tok/s (≤93k ctx), ~4k tok/s at 155k
VRAM ~19GB weights + ~7GB KV pool + headroom

Also verified end-to-end as the backend of a coding agent
(pi): multi-turn tool loop, file
writes, and shell execution all behave correctly.

Known caveat

vLLM logs at load time:

In NVFP4 linear, the weight global scale is different for parallel layers
(e.g. q_proj, k_proj, v_proj).

This checkpoint stores an independent NVFP4 global scale per linear layer;
vLLM reconciles them when fusing q/k/v (and gate/up) GEMMs, which adds a
small one-time rounding cost on those layers. The functional testing above
(tool calling, deep-context retrieval, agent use) shows no observable
degradation, but exact-benchmark users should be aware. If future
llmcompressor releases expose fused-layer scale sharing, a v2 revision may
be published.

How it was made

One-shot RTN quantization with llmcompressor
0.13.0 (compressed-tensors 0.17.0):

recipe = QuantizationModifier(
    targets="Linear",
    scheme="NVFP4",
    ignore=["lm_head", "re:.*visual.*", "re:.*mtp.*"],
)

Calibration: 20 samples from ultrachat_200k (2048 max sequence length),
used only to set activation global scales — no training involved. The
exact recipe ships in this repo as recipe.yaml; the full quantization
and serving setup lives at
preetpatel/qwen3.8_27B-nvidia5090.

Attribution & license

Apache 2.0, inherited from the base model. All credit for the model
weights to Qwen (base model) and
orcarouter (abliterated fine-tune);
this repo only changes the numeric format. The base model is an
abliterated/uncensored variant intended by its authors for AI red-teaming
and research — apply your own judgment and safety measures downstream.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-19NVFP4 quant of Qwen3.8-27B-Uncensored (llmcompressor 0.13, W4A4, 5090-tested)a517ce34.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration