license: apache-2.0
base_model:
- orcarouter/Qwen3.8-27B-Uncensored-FP8
- z-lab/Qwen3.8-27B-DFlash2
pipeline_tag: image-text-to-text
tags: - radiance
- rocm
- amd
- rdna4
- fp8
- speculative-decoding
- dflash2
- vision
- abliterated
- uncensored
Qwen3.8-27B Uncensored FP8 + DFlash2 — radiance container
orcarouter/Qwen3.8-27B-Uncensored-FP8
— Qwen3.8-27B-FP8 with its refusal direction abliterated — as a single .rad container for the
radiance inference engine (AMD RDNA4, ROCm), with its vision tower and the
z-lab/Qwen3.8-27B-DFlash2 block-diffusion
drafter merged in for speculative decoding.
This model's safety alignment has been removed: it answers requests the original model refuses.
The source model's card describes how it was made, how it was evaluated and what it is for.
| File | qwen3.8-27b-uncensored-fp8.rad — 29.43 GiB |
| Weights | the uncensored FP8 checkpoint as it ships: E4M3 with a scale per 128×128 block; the lm_head made block FP8 by the recipe |
| Speculator | DFlash2, FP8 per 128×128 block from its bf16 release; its vocabulary head as 2-bit codes. It was trained on the original model and drafts for this one too: drafts are verified, so they change the speed and never the output |
| Vision | the 27-block vision tower, bf16: images and video in chat requests |
| Context | 262,144 tokens trained; 200K tested |
Serve
radiance --model qwen3.8-27b-uncensored-fp8.rad --tp 2 --max-model-len 200000 --kv-cache-dtype fp8 \
--max-num-seqs 8 --host 0.0.0.0 --port 8000
The server speaks the OpenAI API (/v1/chat/completions, /v1/completions), with tool calls and
structured output, and image_url / video_url parts in chat
messages (PNG, JPEG, WebP, GIF, BMP, TIFF, AVIF; MP4, WebM, MKV, MPEG-TS, AVI with H.264, HEVC, VP8,
VP9 or AV1). The drafter's depth is chosen automatically (--num-speculative-tokens N states one,0 turns speculation off).
Built and tested on 2× Radeon AI PRO R9700 (gfx1201).
How this file was made
rad-convert orcarouter/Qwen3.8-27B-Uncensored-FP8 --draft-model z-lab/Qwen3.8-27B-DFlash2 \
--recipe q38-27b-fp8-df2.recipe -o qwen3.8-27b-uncensored-fp8.rad
The recipe (q38-27b-fp8-df2.recipe in this repository); everything it does not name is the
checkpoint's own:
output.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dflash.draft_head.weight rtn codes=u2 zero=u8 group=128 scale=f16
dflash.fc.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dflash.blk.*.attn_q.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dflash.blk.*.attn_k.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dflash.blk.*.attn_v.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dflash.blk.*.attn_output.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dflash.blk.*.ffn_gate_up.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dflash.blk.*.ffn_down.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16