library_name: ninfer
pipeline_tag: image-text-to-text
license: apache-2.0
language:
- en
- zh
url: https://github.com/Wallawalla47/ninfer-custom
base_model: orcarouter/Qwen3.8-27B-Uncensored-NVFP4
base_model_relation: quantized
tags: - Qwen3.8
- abliterated
- uncensored
- nvfp4
- fp4
- w4a4
- 4bit
- fp8
- quantization
- quantized
- blackwell
- compressed-tensors
- ninfer
- qwen3.8-27b
- 4-bit precision
- multimodal
- vision
- reasoning
- thinking
- long-context
- agentic
- mtp
- speculative-decoding
Qwen3.8-27B-Uncensored-NVFP4-NInferV3
A single-file NInfer V3 engine
artifact of
orcarouter/Qwen3.8-27B-Uncensored-NVFP4,
built for an NVIDIA RTX 5090 (sm_120a). No conversion step at load time: pointninfer-serve at the .ninfer file and run.
This repository is the NInfer form of the orcarouter checkpoint above, an abliterated
(refusal-removed) mixed NVFP4/FP8 build of Qwen3.8-27B. Its NVFP4 and FP8 weights are carried over
bit-exactly for every transformer linear, and two speculative-decoding components (DFlash2 draft
model, indexed proposal head) are added so the artifact runs with NInfer's fastest decode paths.
Recommended engine: I recommend running this artifact with my fork of NInfer, which adds much better prefix caching, tool-calling fixes, speed improvements and various other enhancements over upstream.
Disclaimer (inherited from the source model). The orcarouter checkpoint has had its safety
alignment substantially removed by abliteration, so it will comply with requests the original
Qwen3.8-27B would refuse. The source card releases it strictly for legitimate research
(interpretability, AI-safety and refusal-mechanism study, red-teaming, robustness evaluation) and
says you assume full responsibility for how you use it and what it generates; add your own safety
and moderation layers before exposing it to other users. The abliteration was done by orcarouter,
not by me; this repository only changes the storage format. The full source card, including its
disclaimer, is reproduced at the bottom.
What is in the file
| Component | Contents |
|---|---|
text |
Qwen3.5-family hybrid backbone: 64 layers (Gated Delta-Net linear attention + full attention every 4th layer), hidden 5120, vocab 248,320, native context 262,144 |
vision |
qwen3_5_vision tower (depth 27) for image/video input |
mtp |
Qwen3.5 MTP companion (the abliterated head from the source) for MTP speculative decoding |
dflash2 |
5-layer sliding-attention draft model (window 2048, selector rank 256 / top-16) — the --spec dflash2 path |
proposal |
Indexed 131,072-row proposal head gathered from the source output head — enables cheap copy/proposal drafting via --lm-head-draft |
Format census of the 1,072 prepared weight objects: 112 NVFP4 (MLP projections of layers 0–55,
imported bit-exact), 146 FP8 (e4m3fn, row-scale; 144 attention, GDN and last-eight-layer MLP
objects imported bit-exact, plus the re-quantised embedding and output head), 579 BF16, 96 FP32,
55 Q4, 54 Q5, 28 Q8 and 1 Q6 (grouped, FP16 scales), and 1 INT32 proposal index table.
How it was created
The artifact was built with the NInfer V3 converter using the qwen3_8_27b_nvfp4_orcarouter recipe
(tools/convert/official_recipes.py)
on CPU (~12 min), with components text,vision,mtp,dflash2 and --proposal --proposal-rows 131072.qwen3_8_27b_nvfp4-uncensored.ninfer.conversion.json, included here, is the full per-tensor
conversion report. Sources and treatment:
- orcarouter/Qwen3.8-27B-Uncensored-NVFP4 (the base model: an
llmcompressorcompressed-tensorsmixed-precision checkpoint, per its model card) — the dominant source.- MLP projections of layers 0–55 (112 NVFP4 objects) are imported in their encoded NVFP4
form, bit-exact (import_encoded): the packed FP4 weights and their block scales are carried
over unchanged, with no re-quantisation round-trip. At runtime these drive NInfer's W4A4 NVFP4
kernels, which dequantise weights on the fly against the stored scales and quantise
activations in-kernel. - Attention and GDN projections of all layers, and the MLP projections of the last eight
layers (56–63) (144 FP8 objects) are likewise imported in their encoded per-row FP8 form,
bit-exact: the source's E4M3 weights and per-channel scales are carried over unchanged. - Output head and token embedding are BF16 in the source and are quantised to row-scale FP8
(fp8_row_maxabs), the only weights of the text model that are re-quantised. The conversion
check measured the maximum relative dequantisation error against the BF16 source at 0.034 for
the embedding and 0.023 for the output head (bound: 0.07, the FP8 E4M3 half-ULP). - GDN a/b projections are BF16 in the source and stay BF16. Norms, convolutions,
dt_bias/a_logand the other non-projection tensors keep their source BF16/FP32 formats. - Vision tower and MTP companion are taken from the source's BF16 tensors and stored with the
converter's grouped formats (the vision tower in mixed BF16 + Q4/Q5/Q6/Q8; the MTP companion in
BF16 + Q8).
- MLP projections of layers 0–55 (112 NVFP4 objects) are imported in their encoded NVFP4
- DFlash2 draft — the 5-layer sliding-attention draft model is grafted verbatim from the
local DFlash2 source used for my other NInfer V3 artifacts, so the--spec dflash2path runs
on this backbone. It was not retrained for the abliterated weights. - Indexed proposal head — 131,072 rows gathered from the source output head, ranked by a
frequency corpus, stored as grouped Q4 plus an INT32 token-id table; this is what--lm-head-draftserves copy-heavy proposals from. - Resources — the six tokenizer / chat-template / generation-config / preprocessor resources
are byte-identical to the official NInfer artifact's, not the orcarouter files (the chat
template is the official NInfer one, derived from Qwen's;generation_config.jsonand the
preprocessor configs carry the same values as the source).
Verification. The converter's own checks passed for every quantisation family, and an
exhaustive replay of the deterministic conversion pipeline from the source files (2026-09-27)
byte-compared all 1,072 weight objects plus the auxiliary and resource objects against the artifact:
zero differences. Beyond that I did not measure quality (no perplexity or benchmark numbers for
this artifact), and the source card also lists the quantised checkpoint's capability retention as
not yet measured.
How to run
Requires the NInfer V3 engine and an NVIDIA RTX 50-series GPU (built and tuned on an RTX 5090; the
weights take ~21.4 GiB of VRAM). I recommend my fork
of NInfer, but it is not required. The launch command below uses options from the fork; if you run
upstream NInfer, drop any option its ninfer-serve --help does not list. Prebuilt Windows
executables of the fork are on its releases page and need an NVIDIA driver of 580 or
later; on other systems, build from source (quick start).
set NINFER_CUDA_SYNC=blocking
ninfer-serve.exe qwen3_8_27b_nvfp4-uncensored.ninfer --host 127.0.0.1 --port 8080 --max-context 220000 --max-concurrency 2 --spec dflash2 --draft-tokens 7 --lm-head-draft --ngram-draft-tokens 15 --ngram-min-match 12 --kv-dtype int8 --preserve-thinking --host-cache-mib 52000 --pending-timeout-ms 900000 --prefill-chunk 4096 --kv-capacity auto --vram-headroom-mib 0 --log-colours on --ngram-archive-mib 4096 --ngram-session-mib 256 --ngram-native-sessions --request-log-jsonl log.json --default-thinking-budget 16384 --thinking-budget-message "Considering the limited time available to the user, I must stop thinking now. Time to act:" --tolerant-tool-calls --prefix-cache-file prefixes-uncensored.cache --vision --vision-offload on
This is the launch I use for the NVIDIA NVFP4 artifact, with one change: --max-concurrency 2
instead of 4. This artifact's weights are about 0.9 GiB larger (21.4 GiB against 20.5 GiB), and on
a 32 GB RTX 5090 with --max-concurrency 4 or 3 startup fails because the engine's minimum runtime
reservation no longer fits after the weights. With 2 it starts with a 221,952-token int8 KV cache
(--kv-capacity auto sizes the KV pool to the free VRAM) and about 95 MiB of VRAM left over, so
if other programs use more VRAM on your machine, lower --max-context. Stop any other resident
model first.
The command runs DFlash2 speculative decoding with 7 draft tokens, the indexed proposal head
(--lm-head-draft) and n-gram copy drafting, int8 KV, the hybrid prefix cache (the default) with a
52,000 MiB tier in system RAM (--host-cache-mib) that is saved to and restored from--prefix-cache-file across restarts, a thinking budget of 16,384 tokens, and recovery of
slightly malformed tool calls (--tolerant-tool-calls). --vision enables image and video input,
and --vision-offload on keeps the vision tower in system RAM instead of VRAM, which frees about
0.8 GiB of VRAM for the KV cache. NINFER_CUDA_SYNC=blocking stops the server keeping one CPU core
at 100% while it runs, at a measured cost of about 1% in output tokens per second.ninfer-serve.exe --help lists every option.
Tested on 2026-10-01 on an RTX 5090 (Windows) with the fork's Windows build: the command above
started with this artifact and served a thinking chat request, a tool call, an image request
(which answered "Red and blue" for a red image with a blue stripe) and a repeated prompt that hit the
prefix cache. In that handful of requests decode ran at about 265–295 tokens per second, with
DFlash2 accepting about 45–53% of drafted tokens; this is a smoke test, not a benchmark, and I have
not measured speed or quality against the other artifacts.
Original model card (copied verbatim from orcarouter/Qwen3.8-27B-Uncensored-NVFP4)
The following is the README of the source checkpoint
orcarouter/Qwen3.8-27B-Uncensored-NVFP4,
copied from its source repository for reference.
license: apache-2.0
base_model: Qwen/Qwen3.8-27B
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: transformers
language:
- en
- zh
tags: - abliterated
- qwen
- qwen3
- qwen3.8
- uncensored
- ai-red-team
- red-teaming
- nvfp4
- fp4
- fp8
- mixed-precision
- compressed-tensors
- vllm
- vision-language
- function-calling
- reasoning
- mtp
Qwen3.8-27B-Uncensored-NVFP4
An abliterated (refusal-removed) & dynamic mixed-precision NVFP4 + FP8 build of Qwen's Qwen3.8-27B — for Blackwell FP4
One Gateway. Every Model. — Route Smarter · Ship Safer · Spend Less.
Run via API · API Endpoint · Website · Model Catalog · Model Card · GitHub · Discord · X
An abliterated (refusal-removed) and dynamic mixed-precision NVFP4 + FP8 quantized
build ofQwen/Qwen3.8-27B— a 27B-parameter dense,
hybrid-attention (Gated DeltaNet linear + full attention) native vision-language model with
flexible thinking control, tool-calling, and an MTP speculative-decoding head. This build removes
the safety refusal direction, then quantizes the bulk feed-forward layers to 4-bit NVFP4 while
keeping precision-sensitive layers and the KV cache at FP8, so accuracy is better preserved
than uniform W4A4. 262K context, tools + reasoning + MTP + vision preserved.
Browse all models in the OrcaRouter Model Catalog.Sibling releases: •
Qwen3.8-27B-Uncensored— BF16 source •Qwen3.8-27B-Uncensored-FP8— block-FP8 for vLLM •Qwen3.8-27B-Uncensored-GGUF— 2-bit→16-bit for llama.cpp •Qwen3.8-27B-Uncensored-MLX— MLX for Apple Silicon (2 / 4 / 8-bit).
⚠️ Disclaimer — read before use
This model has had its safety alignment substantially removed via abliteration
(orthogonalizing the refusal direction out of the residual stream). As a direct consequence:
- It will comply with harmful, unethical, offensive, or illegal requests that the
originalQwen3.8-27Bwould refuse. It has no meaningful built-in guardrails. - It is released strictly for legitimate research — interpretability, AI-safety and
refusal-mechanism study, red-teaming, robustness evaluation, and controlled experiments. - You assume full responsibility and liability for how you use it and for everything it
generates. Do not deploy it to end users or in production without adding your own safety,
moderation, and abuse-prevention layers. - Use must comply with the Apache 2.0 License
inherited from the base model, and all laws and regulations that apply to you. - The authors and uploaders accept no liability for any misuse or harm arising from this
model. Its outputs do not reflect the views of the uploaders or of Qwen / Alibaba.
By downloading or using this model you acknowledge and accept the above.
Model details
| Base model | Qwen/Qwen3.8-27B |
| Architecture | Qwen3_5ForConditionalGeneration — 64 layers, hidden 5120, hybrid Gated DeltaNet (48 linear-attention + 16 full-attention, interval 4), native VL tower + MTP head |
| Modification | Abliteration (refusal-direction removal) then dynamic mixed-precision NVFP4 + FP8 quantization |
| Quantization | Mixed-precision (compressed-tensors) — NVFP4 (W4A4) on the FFN of layers 0–55, FP8 (W8A8 dynamic) on attention / GDN projections / the last 8 layers' FFN / lm_head, and a static FP8 KV cache |
| Format | safetensors, resharded to ≤ 5 GB shards (5 + 1 shards, 23.4 GB, 1968 tensors) |
| Precision | FP4 E2M1 group-16 (+ FP8-E4M3 block scale + FP32 global scale) and FP8 E4M3; vision tower / norms / GDN in_proj_a/b / embeddings / MTP head kept in BF16 |
| Preserved | Full vision-language tower and MTP speculative-decoding head (drop-in for the base) |
| Context | 262,144 tokens |
Abliteration
Refusal-direction removal following Arditi et al. (2024), Refusal in Language Models Is
Mediated by a Single Direction. A single refusal direction r (k = 1) is estimated as the
massive-activation–masked mean-difference of harmful − harmless last-token residuals at
layer 38 (round(0.6 × 64)), on AdvBench (harmful) vs Alpaca (harmless). r is then
orthogonalized out of every residual-writing matrix — W' = W − r(rᵀW) — computed in
float32:
| Component | matrices edited |
|---|---|
self_attn.o_proj (16 full-attention layers + MTP) |
17 |
linear_attn.out_proj (48 linear-attention / GDN layers) |
48 |
mlp.down_proj (64 layers + MTP) |
65 |
embed_tokens (row space) |
1 |
| Total | 131 |
The vision tower is untouched and the MTP head is abliterated consistently with the
main model, so speculative decoding keeps working. This is the same abliterated BF16 base as theFP8,GGUF andBF16 releases — only the
quantization differs.
Dynamic mixed-precision (NVFP4 + FP8) scheme
Rather than quantizing every linear uniformly, precision-sensitive layers are kept at FP8 while
only the bulk feed-forward layers go to 4-bit:
| Component | Precision | # linears |
|---|---|---|
MLP gate/up/down, layers 0–55 |
NVFP4 — FP4 E2M1, group 16, FP8-E4M3 block scale + FP32 global scale (W4A4) | 168 |
self_attn.{q,k,v,o}_proj, GDN in_proj_qkv/in_proj_z/out_proj |
FP8 — W8A8, per-channel weight, dynamic per-token activation | 208 |
MLP gate/up/down of the last 8 layers (56–63), lm_head |
FP8 (W8A8 dynamic) | 25 |
| KV cache | FP8 — static, per-tensor | — |
Vision tower, GDN in_proj_a/in_proj_b, all norms, embeddings, MTP head |
BF16 (unquantized) | 206 |
- Weights: round-to-nearest. NVFP4 packs FP4 (E2M1) in groups of 16 with an FP8-E4M3 block
scale and an FP32 per-tensor global scale; FP8 uses per-output-channel scales. - Activations: NVFP4 layers use dynamic per-token FP4 with a calibrated global scale; FP8
layers use dynamic per-token FP8 (no static activation scale). - KV cache: static per-tensor FP8, calibrated.
- Calibration: 512 samples — 75%
tatsu-lab/alpaca- 25% compliant harmful completions,
enable_thinking=False, sequence length 2048 — used only
for the NVFP4 activation global scales and the static FP8 KV-cache scales, keeping calibration on
the activation distribution the abliterated model actually produces.
- 25% compliant harmful completions,
- Built with
llmcompressor(QuantizationModifier, two config groups +kv_cache_scheme),format: mixed-precision. Split: 168 NVFP4 / 233 FP8 / 206 BF16 linears.
vLLM serves this through the compressed-tensors path: the FP4 layers use FP4 tensor cores on
Blackwell, while the FP8 layers run on Hopper-class and newer.
Intended use
- Research into refusal mechanisms, alignment, and interpretability.
- Red-teaming and safety / robustness evaluation in controlled environments.
- Uncensored generation for authorized, lawful research settings.
Out of scope
- Any use that violates the base model's Apache 2.0 license or applicable law.
- Deployment to the public or to end users without additional safety and moderation layers.
- Generating content intended to harm, harass, defraud, or endanger people.
Evaluation
Abliteration is a weight edit shared across all releases of this model, so the refusal
behavior of this checkpoint tracks the BF16 / FP8 builds. On the byte-identical-schemeFP8 build, harmful-prompt
refusal collapses from 64–99% (base) to 0–6% (thinking off) and ≤ 1.7% (thinking on),
while benign over-refusal drops (XSTest-safe 5.6% → 0.4%) and capability stays within ±1.3 pts
of the base (MMLU 84.3 → 84.7, MMLU-Pro 77.6 → 76.8, GSM8K 90.0 → 88.7, CMMLU 81.4 → 80.8). See
that model card for the full tables.
Quant-specific numbers pending. Capability-retention and perplexity for this
NVFP4 + FP8 mixed-precision checkpoint have not yet been measured — the FP4 layers require
Blackwell FP4 tensor cores to run natively, and this build is released for evaluation on
that hardware. Numbers will be added here once benchmarked. As a mixed 4-bit/8-bit checkpoint it
is expected to trade a little accuracy for size versus the FP8 build; the dynamic split (only the
less-sensitive FFN layers at FP4) is designed to keep that loss small.
Multimodal (vision)
The vision tower is preserved — all 167 visual.* weight tensors are kept in BF16 and the
merger / image + video preprocessor configs are intact, so this stays a full vision-language model
(Qwen3_5ForConditionalGeneration), a drop-in for the base. Abliteration only edits the
language-model residual writers, so image understanding is architecturally unaffected (and
image-conditioned refusals are reduced along with text ones). Serve without--language-model-only to use vision.
Usage
Self-host with vLLM (OpenAI-compatible)
Requires a recent vLLM (≥ 0.27, with compressed-tensors). The FP4 layers need a Blackwell
GPU (B200 / GB200 / RTX 50-series) for native FP4 tensor cores; the FP8 layers run on Hopper-class
and newer.
docker run -d --name qwen38-uncensored-nvfp4 --gpus all --ipc=host --shm-size=8g \
-v /path/to/Qwen3.8-27B-Uncensored-NVFP4:/model:ro \
-p 8000:8000 vllm/vllm-openai:v0.27.1 \
--model /model --served-model-name Qwen3.8-27B-Uncensored \
--speculative-config '{"method":"mtp","num_speculative_tokens":2}' \
--gpu-memory-utilization 0.9 \
--max-model-len 262144 --max-num-seqs 96 \
--trust-remote-code \
--reasoning-parser qwen3 \
--enable-auto-tool-choice --tool-call-parser qwen3_coder
The mixed-precision
quantization_config(including the FP8kv_cache_scheme) is read fromconfig.json— do not pass--quantizationor--kv-cache-dtype.--speculative-config mtp
enables the preserved MTP draft head.
Reasoning (thinking) toggle
Thinking is on by default (Qwen3.8). Toggle it per request via chat_template_kwargs; the
reasoning trace is returned in the reasoning field (--reasoning-parser qwen3).
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
resp = client.chat.completions.create(
model="Qwen3.8-27B-Uncensored",
messages=[{"role": "user", "content": "Prove that sqrt(2) is irrational."}],
extra_body={"chat_template_kwargs": {"enable_thinking": True}},
)
print(resp.choices[0].message.reasoning) # thinking trace
print(resp.choices[0].message.content) # final answer
Tool calling
Standard OpenAI tools + assistant tool_calls + role: tool result messages are supported,
including multi-turn (feed the tool result back for a follow-up answer). Parsed by--tool-call-parser qwen3_coder.
Via OrcaRouter (hosted API — no setup)
Served on OrcaRouter through the OpenAI-compatible
gateway (262K context, tools + reasoning). Grab an API key at
orcarouter.ai (sk-orca-...).
from openai import OpenAI
client = OpenAI(base_url="https://api.orcarouter.ai/v1", api_key="sk-orca-...")
resp = client.chat.completions.create(
model="qwen/qwen3.8-27b",
messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)
Hardware requirements & performance
Software
- vLLM ≥ 0.27 with
compressed-tensors(Qwen3.5/3.8 support) — e.g.vllm/vllm-openai:v0.27.1.
Compute
- The NVFP4 (FP4) layers require a Blackwell GPU (B200 / GB200 / RTX 50-series) for native FP4
tensor cores. The FP8 layers run on Hopper (H100 / H200) and newer.
Memory
- Weights: ~23 GB (mixed 4-bit / 8-bit), vs the ~56 GB BF16 checkpoint.
- Minimum ~32 GB VRAM for weights + a small KV cache; the full 262K context needs substantial
extra KV cache (the checkpoint already stores the KV cache in FP8). - Recommended: a single Blackwell B200 (or larger) for the full FP4 path.
Throughput / concurrency
- Continuous batching; concurrency bounded by
--max-num-seqsand the KV cache that fits after
weights are loaded. The MTP draft head gives a large decode speedup on real workloads.
Bias, risks, and limitations
- Safety guardrails removed — the model will produce harmful, biased, or offensive content
on request. See the disclaimer above. - It inherits any biases and limitations of the base
Qwen3.8-27B. - Mixed NVFP4 + FP8 is not lossless versus BF16; quant-specific capability impact for this
build is not yet measured (see Evaluation). - The FP4 layers require Blackwell hardware to run natively; on pre-Blackwell GPUs the FP4 path is
unavailable.
License
Apache 2.0, inherited from the base modelQwen/Qwen3.8-27B. Abliteration and quantization do
not change the underlying license obligations.
Changelog
2026-08-21 — Fixed vLLM loading error (lm_head.weight_scale)
Earlier revisions failed to load in vLLM with:
ValueError: There is no module or parameter named 'lm_head.weight_scale' in Qwen3_5ForCausalLM.
The available parameters belonging to lm_head (ParallelLMHead) are: {'lm_head.weight'}
Cause: the output head (lm_head) had been quantized to FP8, so the checkpoint shipped alm_head.weight_scale tensor. vLLM's Qwen3_5ForCausalLM always builds lm_head as an
unquantized ParallelLMHead (only a weight parameter), leaving the extra scale with no
destination and aborting the load.
Fix: lm_head is now kept unquantized — restored to the original BF16 weight, lm_head.weight_scale
removed, and lm_head moved to the quantization ignore list (matching the INT8 build). Onlyconfig.json, model.safetensors.index.json, and model-00005-of-00005.safetensors changed; all
other tensors (FP4 body, FP8 layers, FP8 KV scales, MTP head) are byte-for-byte identical.