license: other
license_name: qwen-community-1.0
license_link: LICENSE
base_model: orcarouter/Qwen3.8-Flash-Next-Uncensored
base_model_relation: quantized
pipeline_tag: image-text-to-text
tags:
- nvfp4
- fp8
- modelopt
- sglang
- blackwell
- uncensored
- abliterated
Qwen3.8-Flash-Next-Uncensored-NVFP4 (beginning-ai)
A mixed-precision quantization of
orcarouter/Qwen3.8-Flash-Next-Uncensored
(revision e096800036ec20da7e2442dcd4044a004d4e99fa), made to run on one 96 GB
Blackwell GPU with SGLang. That model is an abliterated (refusal-removed) BF16 build
of Qwen/Qwen3.8-Flash-Next.
This checkpoint has the same structure, formats and file layout as our censored
quant, beginning-ai/Qwen3.8-Flash-Next-NVFP4.
The two can be swapped by changing only the model path. Its activation scales were
measured again on this model.
- 120 GiB on disk. The BF16 source is 335 GiB.
- About 78 GB of weights on the GPU. The 47.7 GiB n-gram PLE table stays in host
RAM or on NVMe. - On one RTX PRO 6000, single-stream decode runs at about 265 tok/s at 2K
context and about 240 tok/s at 200K, with a 786K-token KV cache and 524K
context. - Nothing is removed: the vision tower and the multi-token-prediction (MTP)
draft layer are kept.
Warning. The source model has had its refusal behavior removed. It will follow
harmful, unethical or illegal requests that the original model refuses, and it has
no built-in guardrails. Use it only where you add your own safeguards, and follow
the license and the laws that apply to you. We did not change or measure its
refusal behavior; see the source model card.
How this differs from the source and the base
From Qwen's base to orcarouter's source. We compared every tensor with the
official base. orcarouter changed 149 of the 1,658 tensors, all of them weights
that write into the residual stream:
- the output projections of all 36 linear-attention layers and 12 sparse-attention layers;
- the routed-expert and shared-expert down projections in all 48 layers;
- the token embeddings and one PLE projection;
- 3 tensors in the MTP layer.
The router, the norms, the vision tower, the output head and the 47.7 GiB n-gram
table are byte-identical to the base.
From orcarouter's source to this checkpoint:
| Part | Source | This checkpoint | Size here |
|---|---|---|---|
Routed experts (48 layers × 512) mlp.experts.*.{gate,up,down}_proj |
BF16 | NVFP4 (4-bit weights and activations, group 16) | 68.0 GiB |
Shared experts mlp.shared_expert.* |
BF16 | NVFP4 weights, BF16 activations (W4A16_NVFP4, group 16) |
0.13 GiB |
Linear-attention projections (36 layers) linear_attn.{in_proj_qkv,in_proj_z,out_proj} |
BF16 | FP8, 128×128 block scales, dynamic activations (FP8_PB_WO) |
1.93 GiB |
Sparse-attention projections (12 layers) self_attn.{q,k,v,o}_proj |
BF16 | FP8, 128×128 blocks (FP8_PB_WO) |
0.65 GiB |
MTP draft layer routed experts mtp.layers.0.mlp.experts |
BF16 | NVFP4 (group 16) | 1.49 GiB (whole MTP layer) |
| N-gram PLE table | BF16 | FP8, copied byte-for-byte from Qwen/Qwen3.8-Flash-Next-FP8 (the table is unchanged in the source) | 47.7 GiB |
| Output head, embeddings, router, shared-expert gate, norms, linear-attention state parameters, hyper-connections, rest of the MTP layer, vision tower | BF16 | BF16, unchanged (byte-for-byte) | 4.6 GiB (plus the MTP layer's BF16 part) |
The full per-layer list is in hf_quant_config.json (quant_algo: MIXED_PRECISION).
Byte counts and the calibration report are in MANIFEST.json.
How it was made
- Quantizer. NVIDIA ModelOpt 0.47.0 tensor quantizers (
NVFP4QTensor,FP8QTensor), applied to orcarouter's BF16 weights. - Shared scale for gate and up. SGLang runs each expert's
gate_projandup_projas one fused GEMM with one global scale. So both halves are quantized
against one sharedweight_scale_2, as ModelOpt does for fused layers. A build
check confirms that all 24,624 fused pairs, plus the 512 MTP experts, have equal
scales. - Activation scales, measured on this model. We served a first build of this
model in SGLang with a forward hook on every routed-MoE layer. The hook recorded the
MoE inputs and routing for 16 real coding-agent prompts (332,244 tokens) and saved
100,123 rows per layer. From that capture:gate_proj/up_proj: one scale per layer, from the largest MoE input over
every token.down_proj: one scale per expert, from that expert's own activations
(silu(gate) * up), computed with the BF16 weights on the tokens routed to it.
24,184 of the 24,576 experts had at least 32 routed tokens; the median is about
1,350. The rest take the larger of their own maximum and the layer's 95th
percentile.- MTP experts: one scale per layer, from the same capture.
- No scales were taken from another release.
- No quantization of the output head. It stays BF16, as in NVIDIA's and
RadixArk's releases of the base. A 4-bit head changed the top token at about 20%
of positions in our tests.
Results
Test machine
| GPU | 1 × NVIDIA RTX PRO 6000 Blackwell Workstation Edition (96 GB, SM120), driver 580 |
| CPU | AMD Ryzen 9 9950X3D (16 cores) |
| Host RAM | 64 GB. The PLE table does not fit, so it streams from NVMe. |
| Storage | 1 × PCIe Gen4 NVMe (2 TB) for the weights, the PLE table and the KV disk cache |
| Software | beginning-ai/sglang tag qwen38-fn-nvfp4-2026-10-10 (see Serving); FP8 KV cache; MTP speculative decoding (3 steps, 4 draft tokens) with rejection sampling and a 64K-token draft vocabulary |
Correctness checks
| Check | This checkpoint | Our censored quant |
|---|---|---|
| Conversation replay: real DeepSWE agent turns at 8K–200K context that produce a valid tool call | 120 of 120 | 120 of 120 |
| Decode compared with prefill: severe differences (decode log-prob > −2 but prefill < −10) | 0 of 5,032 tokens | 0 of 6,357 tokens |
| Sampled tokens outside prefill's top 20 | 0.10% | 0.14% |
Coding agent: DeepSWE smoke test
DeepSWE v1.1 with the mini-swe-agent harness,
reasoning effort medium, 524K context. We ran 4 tasks that our censored quant
solves, one attempt each, to check that the model works as an agent. This is a
smoke test, not a benchmark score.
| Task | Result |
|---|---|
| etree-xml-diff-patch | solved |
| fd-deterministic-multi-key-sorting | solved |
| opa-template-string-reconstruction | solved |
| httpx-deterministic-cookie-store | not solved: 114 of 115 new-feature tests and all 1,281 existing tests passed |
Speed
All numbers are for one stream on the test machine above. The prompts are real
coding-agent prompts. Each run generates 1,024 tokens with temperature 1.0, top_p
0.95 and top_k 20.
| Context | Decode (tok/s) | Our censored quant |
|---|---|---|
| 2K | 266 | 268 |
| 64K | 246 | 262 |
| 200K | 241 | 233–240 |
Decode is the median of 4 seeds at 2K and 64K. At 200K it is the total over 4
prompts × 4 seeds, with a warm page cache.
| Prompt | Time to first token |
|---|---|
| 2K tokens, not cached | 0.18 s |
| 32K tokens, not cached | 2.7 s |
| 128K tokens, not cached | 11.8 s |
| Agent turn: 2K new tokens after a cached 128K prefix | 0.23 s |
| Capacity | |
|---|---|
| KV cache on the GPU (FP8) | 786,432 tokens |
| Context length | 524,288 tokens (YaRN factor 2 over the native 262,144) |
Serving
The results above were measured with this exact SGLang build:
- Commit
93540334d2
(tagqwen38-fn-nvfp4-2026-10-10
in beginning-ai/sglang). - It is upstream SGLang
f3fe753475plus 9
commits
(compare).
The censored model card
lists what each commit adds.
pip install -e "git+https://github.com/beginning-ai/sglang@93540334d2ce687f0677bebb42dbd252b218e92f#egg=sglang&subdirectory=python"
We have not tested this checkpoint on unpatched SGLang.
python3 -m sglang.launch_server \
--model-path beginning-ai/Qwen3.8-Flash-Next-Uncensored-NVFP4 \
--trust-remote-code \
--quantization modelopt_fp4 \
--dtype bfloat16 \
--kv-cache-dtype fp8_e4m3 \
--page-size 64 \
--context-length 262144 \
--linear-attn-decode-backend flashinfer \
--linear-attn-prefill-backend flashinfer \
--mamba-radix-cache-strategy extra_buffer \
--speculative-algorithm NEXTN \
--speculative-num-steps 3 \
--speculative-eagle-topk 1 \
--speculative-num-draft-tokens 4 \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder
Notes:
- PLE table on NVMe. If the 47.7 GiB PLE table does not fit in host RAM, as on
our 64 GB machine, add--ple-offload-backend nvme(patched SGLang). - 524K context. Add
--context-length 524288and setSGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1. Also pass--json-model-override-argswith a YaRNrope_parametersblock:rope_type: "yarn",factor: 2.0,original_max_position_embeddings: 262144, and the base
config'smrope_*/rope_theta/partial_rotary_factorvalues. - Decode kernels.
SGLANG_FP8_ROW_WEIGHTS=1(FP8 output head and FP8 draft
projections, converted at load time) andSGLANG_NVFP4_DECODE_MOE=1turn on the
patched decode kernels.
Files
| File | Content |
|---|---|
layer-*-experts-*.safetensors |
Routed experts (NVFP4) |
model-quant-00000.safetensors |
FP8 attention projections, NVFP4 shared experts |
model-mtp-experts-nvfp4.safetensors |
MTP draft-layer routed experts (NVFP4) |
model-ple-*.safetensors |
N-gram PLE table (FP8, from Qwen's FP8 release) |
model-bf16-*.safetensors |
All BF16 tensors |
MANIFEST.json |
Provenance, byte accounting, calibration report |
VALIDATION.txt |
Layout and fused-scale checks run at build time |
License and credits
- License: the Qwen Community License 1.0, the same as the base model (see
LICENSE). - Source model: orcarouter/Qwen3.8-Flash-Next-Uncensored
by OrcaRouter. - Base model: Qwen/Qwen3.8-Flash-Next
by the Qwen team. The PLE table is from
Qwen/Qwen3.8-Flash-Next-FP8. - Quantization: NVIDIA TensorRT Model Optimizer (ModelOpt).