license: other
license_name: qwen-community-1.0
license_link: LICENSE
base_model:
- orcarouter/Qwen3.8-Flash-Next-Uncensored
pipeline_tag: image-text-to-text
tags: - radiance
- rocm
- amd
- rdna4
- moe
- 4-bit
- int8
- gptq
- speculative-decoding
- mtp
- vision
- abliterated
- uncensored
Qwen3.8-Flash-Next Uncensored, 4-bit codebook experts + int8 trunk — radiance container
orcarouter/Qwen3.8-Flash-Next-Uncensored
— Qwen3.8-Flash-Next with its refusal direction abliterated — quantised from its bf16 checkpoint into
a single .rad container for the radiance inference engine (AMD RDNA4, ROCm), with the model's
MTP head for speculative decoding and its vision tower. It is the same recipe as
StillDeadcode/qwen3.8-next-flash-fp8-iq4r-moe.
This model's safety alignment has been removed: it answers requests the original model refuses.
The source model's card describes how it was made, how it was evaluated and what it is for.
| File | qwen3.8-next-flash-uncensored-fp8-iq4r-moe.rad — 113.6 GiB |
| Routed experts | 4-bit codes into a 16-level non-uniform codebook after a 128-point Walsh–Hadamard rotation, one E4M3 scale per 64 weights; codes chosen by GPTQ (10M-token calibration, per-expert down-projection Hessians) and refined by three sweeps of coordinate descent |
| Protected experts | ten experts that carry most of their layer's down-projection energy, kept bf16 |
| Trunk | attention, delta-net and shared-expert linears and the lm_head int8 (W8A8), a scale per 128 columns searched for least error; hyper-connection mixing matrices E4M3 |
| Speculator | the model's MTP head (depth 3 by default) |
| Vision | the vision tower, bf16: images and video in chat requests |
| Context | 262,144 tokens trained; served at 200K |
Quality
The calibration data was recorded on the original model, and this file was not measured against
its own bf16 weights. For the recipe's error, the same recipe on the original model, as
full-vocabulary KL divergence against its bf16, teacher-forced over ~75K positions of chat,
tool-use and code transcripts:
| mean KL | median | p99 | p99.9 | top-1 agreement | |
|---|---|---|---|---|---|
| all tokens | 0.0789 | 0.0037 | 1.65 | 5.33 | 92.85% |
| text and assistant turns | 0.0239 | 0.0029 | 0.32 | 1.03 | 94.57% |
| reference: a second bf16 implementation, all tokens | 0.0399 | 0.0013 | 0.82 | 3.43 | 95.13% |
Serve
The routed experts do not fit two 32 GB cards; the engine keeps the hottest in VRAM and streams the
rest from a pinned host pool, so the host needs about 32 GB of free RAM.
radiance --model qwen3.8-next-flash-uncensored-fp8-iq4r-moe.rad --tp 2 --tp-wire wht6 \
--max-model-len 200000 --max-num-seqs 8 --kv-cache-dtype fp8 \
--placement expert_tiered --host-pool-mib 12288 --gpu-headroom-mib 96 \
--expert-vs-cache-ratio 0.82 --num-speculative-tokens 3 --max-num-batched-tokens 2048 \
--host 0.0.0.0 --port 8000
The recipe is served on 2× Radeon AI PRO R9700 (gfx1201). This file was converted there with no
errors and uploaded without being served first.
The server speaks the OpenAI API (/v1/chat/completions, /v1/completions), with tool calls,
structured output, and image_url / video parts in chat messages.
How this file was made
CALIB=calib/w4nl-calib rad-convert orcarouter/Qwen3.8-Flash-Next-Uncensored \
--tokenizer orcarouter/Qwen3.8-Flash-Next-Uncensored/tokenizer.json \
--recipe qwen4exp-w4nl64-i8-hc8m.recipe -o qwen3.8-next-flash-uncensored-fp8-iq4r-moe.rad
The recipe (qwen4exp-w4nl64-i8-hc8m.recipe in this repository; $CALIB is its calibration data,
which is not published). The first rule that matches a weight decides it, and everything no rule
names is the checkpoint's bf16:
blk.34.ffn_*_exps.407.weight cast dtype=bf16
blk.34.ffn_*_exps.496.weight cast dtype=bf16
blk.44.ffn_*_exps.292.weight cast dtype=bf16
blk.44.ffn_*_exps.350.weight cast dtype=bf16
blk.46.ffn_*_exps.290.weight cast dtype=bf16
blk.46.ffn_*_exps.392.weight cast dtype=bf16
blk.47.ffn_*_exps.122.weight cast dtype=bf16
blk.47.ffn_*_exps.143.weight cast dtype=bf16
blk.47.ffn_*_exps.399.weight cast dtype=bf16
blk.47.ffn_*_exps.445.weight cast dtype=bf16
blk.*.ffn_*_exps.*.weight gptq table=w4nl group=64 scale=fp8_e4m3 scale2=f32 block2=*x* scale2_value=0.0001220703125 transform=fwht128 rule=search cd=3 calib=$CALIB
*_hc_down.weight rtn codes=fp8_e4m3 group=128 scale=f32
*_hc_up.weight rtn codes=fp8_e4m3 group=80 scale=f32
mtp.fc_hidden.weight rtn codes=fp8_e4m3 group=128 scale=f32
mtp.fc_embedding.weight rtn codes=fp8_e4m3 group=128 scale=f32
blk.*.attn_qg.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
blk.*.attn_k.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
blk.*.attn_v.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
blk.*.attn_output.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
blk.*.ssm_inz.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
blk.*.ssm_out.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
blk.*.ffn_gate_up_shexp.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
blk.*.ffn_down_shexp.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
output.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
blk.*.ple_ngram.weight rtn codes=fp8_e4m3 block=*x* scale=bf16 rule=fixed scale_value=1.99317932128906e-4
mtp.draft_head.weight rtn codes=u2 zero=u8 group=128 scale=f16
License
The Qwen community license of the base model; see LICENSE. The source model's card states
Apache-2.0, but its repository ships this license file, the one the base model carries.