license: other
license_name: qwen-community-1.0
base_model:
- orcarouter/Qwen3.8-Flash-Next-Uncensored
- Qwen/Qwen3.8-Flash-Next
base_model_relation: quantized
pipeline_tag: text-generation
library_name: gguf
tags: - gguf
- rocmfp4
- llama.cpp
- strix-halo
- gfx1151
- rocm
- amd
- ryzen-ai-max
- uncensored
- research
Qwen3.8-Flash-Next-Uncensored — ROCmFP4 STRIX_LEAN GGUF — AMD Ryzen AI Max+ 395 / gfx1151
⚠️ Research artifact. Refusal behaviour has been removed. This does not add capability — it
removes guardrails. Use it deliberately, in a context where that is appropriate, and own the output.
Quantized from the BF16 weights published by
orcarouter/Qwen3.8-Flash-Next-Uncensored
— the abliteration work here is theirs, not mine. Go star their repo.
STRIX_LEAN is my size/speed tier for Strix Halo: the Q4_0_ROCMFP4_STRIX_LEAN recipe — ROCmFP4
weights with Strix attention K/V handling, Q5_K token embeddings, and a Q6_K output head.
Converted to BF16 GGUF and quantized by me from their release. 4.78 bpw, 98.49 GiB.
| tensor group | type |
|---|---|
MoE expert weights (ffn_*_exps) |
TYPE_101 (ROCmFP4, 4.251 bpw) |
shared expert (ffn_*_shexp) |
TYPE_101 |
attention (attn_*) |
half TYPE_100, half TYPE_101 |
per_layer_token_embd.weight (PLE, 51.2B params) |
Q5_1 |
token_embd.weight |
Q5_K |
output.weight (lm head) |
Q6_K |
The size matches my aligned build of the same tier to 0.01 GiB — the abliterated checkpoint is
structurally identical, so the quant recipe transfers exactly.
The Q6_K head
output.weight is Q6_K, never 4-bit. Every sampled token passes through the lm head, so its
quantization error lands directly in the argmax. Verified by exact tensor name after both
quantize and split — output.weight is a substring of attn_output.weight, so a loose check
reports success on a 4-bit head.
⚠ Patched llama.cpp required
Needs PR #27742 merged into the ROCmFPX fork. Stock builds will not load this: both theqwen4exp architecture and the Q4_0_ROCMFP4_* tensor types live in that fork.
-DGGML_HIP=ON -DGPU_TARGETS=gfx1151 -DGGML_NATIVE=ON
Measured — Ryzen AI MAX+ 395, gfx1151, ROCm 7.2.4, full 49/49 offload
- generation: 22.5 tok/s (median of 3, unique prompt each run)
- prompt processing: 222 tok/s
- GPU memory: 63.3 GiB resident — identical to the aligned build
GPU-only, full offload. I do not publish partial-offload speeds.
Long context
This model's native max is 262,144, and it runs there on a 128 GB box:
| context | prompt | pp tok/s | gen tok/s | GTT |
|---|---|---|---|---|
| 131,072 | 111,411 | 185 | 15.33 | 69.1 GiB |
| 262,144 | 8,000 | 307 | 22.48 | 72.0 GiB |
| 262,144 | 200,000 | 128 | 10.46 | 74.9 GiB |
The context window is nearly free — GTT grows only ~4 GiB from 8k to 128k, because Qwen Sparse
Attention caps KV. What you pay for is depth: a 200k-token prompt halves generation. It
degrades smoothly rather than falling off a cliff.
Refusal / quality (counts only)
Aligned build vs this one, same prompts, greedy, same harness:
| split | aligned | this build |
|---|---|---|
| Harmful (24) | 0 comply | 22 comply |
| Harmless (12) | 10 ok | 11 ok |
| Quality (8) | 6/8 | 6/8 — same two failures |
Quality is unchanged to the specific failing question, which is the point: the abliteration
flipped refusal without the quant damaging the model. Prompts and completions are not published.
Files
Sharded to stay under HF's 50 GB limit. Point --model at the first shard.
| file | size |
|---|---|
Qwen3.8-Flash-Next-Uncensored-Q4_0-ROCmFP4-STRIX_LEAN-00001-of-00003.gguf |
41.86 GiB |
Qwen3.8-Flash-Next-Uncensored-Q4_0-ROCmFP4-STRIX_LEAN-00002-of-00003.gguf |
41.62 GiB |
Qwen3.8-Flash-Next-Uncensored-Q4_0-ROCmFP4-STRIX_LEAN-00003-of-00003.gguf |
15.01 GiB |
mmproj-Qwen3.8-Flash-Next-Uncensored-BF16.gguf |
0.85 GiB (vision tower) |
Usage
llama-server \
--model Qwen3.8-Flash-Next-Uncensored-Q4_0-ROCmFP4-STRIX_LEAN-00001-of-00003.gguf \
--mmproj mmproj-Qwen3.8-Flash-Next-Uncensored-BF16.gguf \
--host 127.0.0.1 --port 8080 \
--n-gpu-layers 999 --flash-attn on --fit off \
--ctx-size 131072 --threads 16 --jinja
Do not use --no-mmap. The PLE table is streamed from the file through the page cache; forcing
it into anonymous memory gets the process OOM-killed with nothing in the server log.
Acknowledgements
ROCmFPX — defines the ROCmFP4 tensor formats and carries the qwen4exp support merged from
PR #27742. Every file here was produced with its llama-quantize and runs on its runtime. MIT,
based on upstream llama.cpp.
llama.cpp — ggml-org and contributors — the engine,
GGUF format and conversion tooling this is built on.
AMD ROCm — the compute platform targeted here (ROCm 7.2.4, gfx1151).
orcarouter — published the uncensored BF16 checkpoint
this is built from. The abliteration is their engineering; I only converted and quantized it.
Qwen team — the original base model. See base_model; license qwen-community-1.0.