license: apache-2.0
base_model: orcarouter/Qwen3.8-27B-Uncensored
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: transformers
language:
- en
- zh
tags: - qwen3.8
- nvfp4
- compressed-tensors
- fastllm
- rtx-2080-ti
- turing
- abliterated
- uncensored
- vision-language
- mtp
Qwen3.8-27B-Uncensored-NVFP4-FastLLM
An NVFP4 build of orcarouter/Qwen3.8-27B-Uncensored with every linear layer in NVFP4. It is 20.6 GB, down from 55.6 GB in BF16 and 29 GB in FP8.
I made it to serve the model on two RTX 2080 Ti 22 GB cards with FastLLM. Turing has no FP4 hardware. FastLLM stores the NVFP4 weights as-is and unpacks them to FP16 inside the kernel. Decoding is limited by weight reads, so the smaller file is faster on these cards even with the extra unpacking.
Write-up with the full numbers, the launch command and the debugging notes: NVFP4 Without FP4 Hardware: Qwen3.8-27B on Two RTX 2080 Tis at 151.7 tok/s (中文: 2080 Ti 沒有 FP4 也能跑 NVFP4).
⚠️ Inherited disclaimer
The base model has had its safety alignment largely removed by abliteration. Quantizing it does not change that. Everything in the base model's disclaimer applies here. The model will comply with harmful requests. It is meant for research and controlled use, you are responsible for how you use it, and you should add your own moderation before exposing it to anyone else.
What's in it
| Weights | Format |
|---|---|
All linear layers: attention q/k/v/o, GDN projections (in_proj_qkv/z/a/b, out_proj), MLP |
NVFP4 (E2M1 values, one FP8-E4M3 scale per 16 values, one FP32 global scale per tensor) |
lm_head, embed_tokens, vision tower, conv1d, MTP head |
BF16, unchanged |
- The layout matches lyf/Qwen3.8-27B-Heretic-ARA-NVFP4-MTP-VL exactly: tensor names, dtypes, shapes, shard map,
quantization_configandhf_quant_config.json(2,687 tensors,total_size20,558,935,392). Only the base model differs. - As in ModelOpt's recipe, q/k/v share one global scale per layer, the four GDN input projections share one, and MLP gate/up share one.
- No calibration:
input_global_scaleis set to 1.0. FastLLM on sm_75 computes with FP16 activations and does not read it. If you run this on an engine that does W4A4 with real activation scales, expect worse results than a calibrated checkpoint.
How it was made
I converted it with a numpy-only script that runs on the CPU, no GPU needed: qwen38-nvfp4-convert.py.
hf download orcarouter/Qwen3.8-27B-Uncensored --local-dir ./bf16
hf download lyf/Qwen3.8-27B-Heretic-ARA-NVFP4-MTP-VL --include "*.json" "recipe.yaml" --local-dir ./nvfp4-ref
python3 qwen38-nvfp4-convert.py ./bf16 ./out --reference ./nvfp4-ref
To check the converter, I dequantized one attention, one GDN and one MLP layer of lyf's release to BF16 and quantized them again. The packed weights, block scales and global scales came out byte-identical to lyf's.
Results on two RTX 2080 Ti 22 GB (TP=2, FastLLM + DFlash2)
Same machine, same FastLLM build, temperature 0, 512 tokens, median of 3 runs, tok/s:
| code | math | prose | no speculation (code) | |
|---|---|---|---|---|
| orcarouter FP8 | 116.1 | 135.5 | 74.2 | 32.9 |
| this NVFP4 | 151.7 | 178.8 | 101.9 | 45.1 |
With fp8 KV cache and --max_batch 4: 2 concurrent streams run at 107.8 tok/s each, 4 streams at 74.3 each (about 297 combined).
140-question eval (GSM8K 50, MATH-500 30, MuSR 20, HumanEval+ 40):
| no thinking | thinking (effort low) | |
|---|---|---|
| orcarouter FP8 | 132 | 127 |
| this NVFP4 | 131 | 128 |
Trade-off: prefilling a 54K-token prompt takes 50.6 s instead of 40.5 s. Prefill is compute-bound, so the unpacking is pure overhead there.
I have only tested this with FastLLM on sm_75. I have not tested vLLM, SGLang or Blackwell GPUs.
Serving with FastLLM
Tested on FastLLM at upstream a2bf07fd with d4b04876 reverted and PRs #749 and #756 merged (build steps in the previous post). The draft model is z-lab/Qwen3.8-27B-DFlash2.
export CUDA_VISIBLE_DEVICES=0,1
export FASTLLM_CUDA_GRAPH=0
export FASTLLM_DRAFT_QUANT=nvfp4
export FASTLLM_COOPERATIVE_LONG_PREFILL=1 # from PR #756; leave out on stock FastLLM
ftllm server -p ./Qwen3.8-27B-Uncensored-NVFP4-FastLLM \
--tp 2 --max_batch 4 --chunked_prefill_size 4096 \
--gpu_mem_ratio 0.95 --kv_cache_dtype fp8_e4m3 \
--speculative_algorithm dflash \
--speculative_draft_model_path ./Qwen3.8-27B-DFlash2 \
--speculative_num_draft_tokens 8 \
--prefix_cache true --port 8080
Don't pass --tokens: with it set, FastLLM clamps --max_batch back to 1 for this model.
Credits
- Base model: orcarouter/Qwen3.8-27B-Uncensored, derived from Qwen/Qwen3.8-27B. Apache 2.0.
- Layout reference: lyf/Qwen3.8-27B-Heretic-ARA-NVFP4-MTP-VL.
- Inference engine: FastLLM.