← back to catalog · registered 2026-08-27 15:02

kingjones777/Qwen3.8-Flash-Next-Uncensored-ROCmFP4-STRIX_LEAN-GGUF

kingjones777 Qwen GGUF second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/kingjones777%2FQwen3.8-Flash-Next-Uncensored-ROCmFP4-STRIX_LEAN-GGUF"
Response includes
  • classification m-uncensored
  • files 6
  • hub_downloads_all_time 7,691
  • author_summary 11 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
8K
970 last 30d - stable
Likes
5
Model age
6w ago
created 2026-08-27
Downloads over time
Now7.9K→from970↑717%
02.9K5.8K8.7K970 on Aug 267.9K on Oct 11AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 2K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
other
Tags
gguf rocmfp4 llama.cpp strix-halo gfx1151 rocm amd ryzen-ai-max uncensored research text-generation base_model:Qwen/Qwen3.8-Flash-Next

Related

Total size
98.5 GB
Files
6
Quantizations
2
Registered
2026-08-27 15:02
Last updated on HF
2026-09-17 18:42

Files by quantization

BF16 1 file 866 MB
mmproj-Qwen3.8-Flash-Next-Uncensored-BF16.gguf 866 MB f9d4d77c download
Auxiliary files 5 files 98.5 GB
Qwen3.8-Flash-Next-Uncensored-Q4_0-ROCmFP4-STRIX_LEAN-00001-of-00003.gguf 41.9 GB faa00798 download
Qwen3.8-Flash-Next-Uncensored-Q4_0-ROCmFP4-STRIX_LEAN-00002-of-00003.gguf 41.6 GB bd1fbc53 download
Qwen3.8-Flash-Next-Uncensored-Q4_0-ROCmFP4-STRIX_LEAN-00003-of-00003.gguf 15.0 GB b4f283fe download
README.md 5.26 KB 3454f68b download
.gitattributes 1.89 KB 51293f14 download

README current version from Hugging Face


license: other
license_name: qwen-community-1.0
base_model:

  • orcarouter/Qwen3.8-Flash-Next-Uncensored
  • Qwen/Qwen3.8-Flash-Next
    base_model_relation: quantized
    pipeline_tag: text-generation
    library_name: gguf
    tags:
  • gguf
  • rocmfp4
  • llama.cpp
  • strix-halo
  • gfx1151
  • rocm
  • amd
  • ryzen-ai-max
  • uncensored
  • research

Qwen3.8-Flash-Next-Uncensored — ROCmFP4 STRIX_LEAN GGUF — AMD Ryzen AI Max+ 395 / gfx1151

⚠️ Research artifact. Refusal behaviour has been removed. This does not add capability — it
removes guardrails. Use it deliberately, in a context where that is appropriate, and own the output.

Quantized from the BF16 weights published by
orcarouter/Qwen3.8-Flash-Next-Uncensored
— the abliteration work here is theirs, not mine. Go star their repo.

STRIX_LEAN is my size/speed tier for Strix Halo: the Q4_0_ROCMFP4_STRIX_LEAN recipe — ROCmFP4
weights with Strix attention K/V handling, Q5_K token embeddings, and a Q6_K output head.
Converted to BF16 GGUF and quantized by me from their release. 4.78 bpw, 98.49 GiB.

tensor group type
MoE expert weights (ffn_*_exps) TYPE_101 (ROCmFP4, 4.251 bpw)
shared expert (ffn_*_shexp) TYPE_101
attention (attn_*) half TYPE_100, half TYPE_101
per_layer_token_embd.weight (PLE, 51.2B params) Q5_1
token_embd.weight Q5_K
output.weight (lm head) Q6_K

The size matches my aligned build of the same tier to 0.01 GiB — the abliterated checkpoint is
structurally identical, so the quant recipe transfers exactly.

The Q6_K head

output.weight is Q6_K, never 4-bit. Every sampled token passes through the lm head, so its
quantization error lands directly in the argmax. Verified by exact tensor name after both
quantize and split — output.weight is a substring of attn_output.weight, so a loose check
reports success on a 4-bit head.

⚠ Patched llama.cpp required

Needs PR #27742 merged into the ROCmFPX fork. Stock builds will not load this: both the
qwen4exp architecture and the Q4_0_ROCMFP4_* tensor types live in that fork.

-DGGML_HIP=ON -DGPU_TARGETS=gfx1151 -DGGML_NATIVE=ON

Measured — Ryzen AI MAX+ 395, gfx1151, ROCm 7.2.4, full 49/49 offload

  • generation: 22.5 tok/s (median of 3, unique prompt each run)
  • prompt processing: 222 tok/s
  • GPU memory: 63.3 GiB resident — identical to the aligned build

GPU-only, full offload. I do not publish partial-offload speeds.

Long context

This model's native max is 262,144, and it runs there on a 128 GB box:

context prompt pp tok/s gen tok/s GTT
131,072 111,411 185 15.33 69.1 GiB
262,144 8,000 307 22.48 72.0 GiB
262,144 200,000 128 10.46 74.9 GiB

The context window is nearly free — GTT grows only ~4 GiB from 8k to 128k, because Qwen Sparse
Attention caps KV. What you pay for is depth: a 200k-token prompt halves generation. It
degrades smoothly rather than falling off a cliff.

Refusal / quality (counts only)

Aligned build vs this one, same prompts, greedy, same harness:

split aligned this build
Harmful (24) 0 comply 22 comply
Harmless (12) 10 ok 11 ok
Quality (8) 6/8 6/8 — same two failures

Quality is unchanged to the specific failing question, which is the point: the abliteration
flipped refusal without the quant damaging the model. Prompts and completions are not published.

Files

Sharded to stay under HF's 50 GB limit. Point --model at the first shard.

file size
Qwen3.8-Flash-Next-Uncensored-Q4_0-ROCmFP4-STRIX_LEAN-00001-of-00003.gguf 41.86 GiB
Qwen3.8-Flash-Next-Uncensored-Q4_0-ROCmFP4-STRIX_LEAN-00002-of-00003.gguf 41.62 GiB
Qwen3.8-Flash-Next-Uncensored-Q4_0-ROCmFP4-STRIX_LEAN-00003-of-00003.gguf 15.01 GiB
mmproj-Qwen3.8-Flash-Next-Uncensored-BF16.gguf 0.85 GiB (vision tower)

Usage

llama-server \
  --model Qwen3.8-Flash-Next-Uncensored-Q4_0-ROCmFP4-STRIX_LEAN-00001-of-00003.gguf \
  --mmproj mmproj-Qwen3.8-Flash-Next-Uncensored-BF16.gguf \
  --host 127.0.0.1 --port 8080 \
  --n-gpu-layers 999 --flash-attn on --fit off \
  --ctx-size 131072 --threads 16 --jinja

Do not use --no-mmap. The PLE table is streamed from the file through the page cache; forcing
it into anonymous memory gets the process OOM-killed with nothing in the server log.

Acknowledgements

ROCmFPX — defines the ROCmFP4 tensor formats and carries the qwen4exp support merged from
PR #27742. Every file here was produced with its llama-quantize and runs on its runtime. MIT,
based on upstream llama.cpp.

llama.cpp — ggml-org and contributors — the engine,
GGUF format and conversion tooling this is built on.

AMD ROCm — the compute platform targeted here (ROCm 7.2.4, gfx1151).

orcarouter — published the uncensored BF16 checkpoint
this is built from. The abliteration is their engineering; I only converted and quantized it.

Qwen team — the original base model. See base_model; license qwen-community-1.0.

README history 11 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-09-17docs: measured plain vs draft-mtp (median of 3); recommend --spec-draft-n-max 1a4d9d7b13.1 KB
    Loading...
  2. 2026-09-17docs: draft-mtp verified 2K->128K; honest speed caveat; buildable recipe5b10e2812.5 KB
    Loading...
  3. 2026-09-17MTP section: short-context numbers, name the measured Q8_0 head, keep draft-m...811f74d10.5 KB
    Loading...
  4. 2026-09-16Document working draft-mtp (+27.7%) + bundle graph patch11190f810.1 KB
    Loading...
  5. 2026-09-03docs: warn that ngram-mod at >=64K can wedge the GPU (QSA checkpoint desync);...bc53d608.6 KB
    Loading...
  6. 2026-08-28docs: point runtime instructions at the kingjones30/ROCmFPX fork (verified bu...a7b58a37.6 KB
    Loading...
  7. 2026-08-28Acknowledgements: qwen4exp is not in the ROCmFPX fork; applied via the patch ...024ec736.7 KB
    Loading...
  8. 2026-08-28Build section: exact public clone + build steps, verified from a clean checkout8a4ab886.5 KB
    Loading...
  9. 2026-08-27Correct prompt-processing: fixed-prompt measurement on the published files (2...0fb80195.6 KB
    Loading...
  10. 2026-08-27Credit orcarouter for the BF16 uncensored checkpoint; tag as base_modelecfdce85.3 KB
    Loading...
  11. 2026-08-27Card: ROCmFP4 STRIX_LEAN research build, measured + gatedcd697f44.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration