license: other
license_name: qwen-community-1.0
license_link: https://huggingface.co/Qwen/Qwen3.8-Flash-Next/blob/main/LICENSE
base_model:
- Qwen/Qwen3.8-Flash-Next
base_model_relation: finetune
quantized_by: AtomicChat
pipeline_tag: text-generation
library_name: gguf
tags: - atomic-chat
- qwen
- qwen3.8
- flash-next
- moe
- gguf
- imatrix
- llama.cpp
- lora
- abliteration
- refusal-direction
- not-for-all-audiences
How to Run Qwen3.8-Flash-Next Without Refusals Locally
One refusal direction projected out of Qwen's original weights, shipped two ways: three baked GGUF builds, and a 318 MB LoRA adapter for the Flash-Next quants you already have. The measurements behind this card are public.
- On held-out prompts, refusals fall from 92.6% to 1.2% in English and from 80% to 0% in Russian, and to none of 30 with thinking on. MMLU moves by 0.5 points, inside the noise.
- The 93.9 GB build matches the original's token choice 91.1% of the time: closer to the original than our published 94.5 GB quant of the unmodified model (89.6%).
- Already running our Flash-Next quants? Add the adapter with
--lora-scaledinstead of downloading a new model.
The builds
| Build | In memory | On SSD | Total | Mean KLD | Same top-1 | Refused: EN / RU / XSTest |
|---|---|---|---|---|---|---|
AD-3.86bpw-IQ4_XS |
53.3 GB | 32.0 GB | 85.3 GB | 0.113 | 87.8% | 0% / 0% / 0.4% |
AD-4.25bpw-Q4_K_M |
61.9 GB | 32.0 GB | 93.9 GB | 0.060 | 91.1% | 0% / 0% / 0.8% |
AD-4.85bpw-Q5_K_M |
75.3 GB | 32.0 GB | 107.3 GB | 0.037 | 93.0% | 1.2% / 0% / 0.8% |
AD-4.25bpw-Q4_K_M is the one to take if it fits. AD-3.86bpw is the one
whose in-memory part fits a 64 GB Mac next to the n-gram table on SSD, the same
way our 85 GB build
runs there; this build has not been run on a Mac yet.
KLD and top-1 are against the original, unmodified BF16 model, so they hold
the ablation and the quantization together. For scale: the ablation alone, on
BF16 weights, is 0.020; our plain Q8_0 of the original is 0.017.
Against the other way to get the same model, our original quant of the same
size with the adapter on top, the baked builds are closer to the original. The
gain is the layout, our current recipe, which the original quants predate; the
adapter itself adds about 0.007 to any quant:
| Total | Mean KLD | Same top-1 | |
|---|---|---|---|
baked, this repo: AD-3.86bpw-IQ4_XS |
85.3 GB | 0.113 | 87.8% |
original quant AD-3.84bpw-IQ4_XS + adapter |
84.9 GB | 0.234 | 82.4% |
baked, this repo: AD-4.25bpw-Q4_K_M |
93.9 GB | 0.060 | 91.1% |
original quant AD-4.27bpw-Q4_K_M + adapter |
94.5 GB | 0.090 | 89.2% |
Naming
Files are named by measured bits per weight, with the type tags of our
original quants, so
one tag means one size in both repos: IQ4_XS ~85 GB, Q4_K_M ~94 GB, Q5_K_M
~107-110 GB. The tag is not the experts' type: in AD-4.25bpw-Q4_K_M they are
IQ3_XXS and IQ4_XS, ffn_down_exps IQ4_NL, the n-gram table Q4_1.
What changed, and what did not
Original against ablated, both BF16, on prompts never used to pick anything:
| Original | Ablated | |
|---|---|---|
| JBB harmful behaviours (EN, 81), refused | 92.6% | 1.2% |
| Aya red-teaming (RU, 100), refused | 80.0% | 0% |
| XSTest safe prompts (250), refused | 5.2% | 0.8% |
| Thinking low / xhigh (30 prompts), refused | 97% / 80% | 0% / 0% |
| MMLU, 2000 questions | 87.30% | 86.80% (-0.5, McNemar p = 0.31) |
| Tool calls (20) | 20/20 valid | 20/20 valid |
| Needle at 30k tokens (3 depths) | 3/3 | 3/3 |
| Mean KLD to the original, neutral / code | - | 0.020 / 0.017 |
Read by hand, the answers are complete and in the language asked. Refusal is
counted by the opening of the reply, so it is indicative, not a judge; empty or
degenerate replies count as damage, and there were none.
The adapter
lora/Qwen3.8-Flash-Next-Abliterated-Uncensored-LoRA.gguf: rank 1, 318 MB, 146 sites. It
applies to any GGUF of the original Flash-Next weights, including our three
published quants:
llama-server -m Qwen3.8-Flash-Next-AD-4.27bpw-Q4_K_M-M64-00001-of-00033.gguf \
-ngl 99 --jinja -fit off \
--lora-scaled lora/Qwen3.8-Flash-Next-Abliterated-Uncensored-LoRA.gguf:1.25
Or load it once and set the strength per request:
llama-server -m <model> -ngl 99 --jinja -fit off \
--lora lora/Qwen3.8-Flash-Next-Abliterated-Uncensored-LoRA.gguf --lora-init-without-apply
curl localhost:8080/v1/chat/completions -H 'Content-Type: application/json' -d '{
"messages": [{"role": "user", "content": "..."}],
"lora": [{"id": 0, "scale": 1.25}]
}'
Use scale 1.25. At 1.0, 47% of English validation prompts still refused;
above 1.25 the model moves further from the original with nothing left to
remove. On our published quants:
| Published quant | Mean KLD without / with | Refused with: EN / RU / XSTest |
|---|---|---|
AD-3.84bpw-IQ4_XS |
0.227 / 0.234 | 1.2% / 0% / 0.4% |
AD-4.27bpw-Q4_K_M |
0.083 / 0.090 | 0% / 0% / 0.8% |
AD-5.00bpw-Q5_K_M |
0.083 / 0.089 | 0% / 0% / 0.8% |
Get started
- Atomic Chat:
the easiest path. Open the app, searchAtomicChat/Qwen3.8-Flash-Next-Abliterated-Uncensored-GGUF,
pick a build, hit Use this model. - llama.cpp: download one build and point
-mat its first shard.
hf download AtomicChat/Qwen3.8-Flash-Next-Abliterated-Uncensored-GGUF \
--include "Qwen3.8-Flash-Next-Abliterated-Uncensored-AD-4.25bpw-Q4_K_M/*" --local-dir .
llama-server \
-m Qwen3.8-Flash-Next-Abliterated-Uncensored-AD-4.25bpw-Q4_K_M/Qwen3.8-Flash-Next-Abliterated-Uncensored-AD-4.25bpw-Q4_K_M-00001-of-00036.gguf \
-ngl 99 -c 32768 --jinja -fit off
As with the original quants:
keep mmap on, pass -fit off, and always --jinja for the model's own chat
template. Shard 2 of every build holds nothing but the n-gram table, so it can
stay on SSD. Needs a llama.cpp build with Qwen3.8-Flash-Next support
(PR #27742).
Sampling as Qwen recommends for the original: temperature 1.0, top_p 0.95,
top_k 20 with thinking; temperature 0.7, top_p 0.8, top_k 20,
presence_penalty 1.5 without.
How it was made
- Direction. Difference of means of the residual stream at the last prompt
token, chat template applied, thinking off: 416 harmful against 416 harmless
English prompts. Taken entering block 34 of 48, averaged over the four
hyper-connection streams, with the harmless-mean component removed. - Edit.
W' = W - 1.25 r rᵀWon every matrix that writes into the
residual: 48 attention / DeltaNet outputs, all 512 experts of all 48 MoE
blocks, 48 shared experts, the n-gram value projection and the token
embedding. The router and the hyper-connection weights are untouched. - Choice. 34 variants screened on validation prompts (row, strength, which
writers, one direction per block, English only or English + Russian); the
numbers on this card come from held-out test prompts. - Bake. The same edit applied to the BF16 weights in f32 and rounded once;
every one of the 146 writers checked to keep 0.25 of its original component
alongr, within 3.4e-4. Then quantized with the importance matrix of our
original quants.
One thing that did not work: editing the attention outputs alone, which is all
that Heretic-style tools reach on this architecture, left 96% of English
refusals in place with this direction. The refusal lives in the experts too.
Against other releases
Measured here at Q8_0, against the same BF16 reference and on the same prompts:
| This release (Q8_0 + adapter) | OrcaRouter Uncensored | heretic-2 | |
|---|---|---|---|
| Refused, EN / RU | 0% / 0% | 0% / 1% | 0% / 0% |
| XSTest refused | 1.2% | 0.8% | 0.4% |
| MMLU (2000) | 86.50% | 86.65% | 87.20% |
| Mean KLD to the original | 0.026 | 0.022 | 0.019 |
| Same top-1 | 94.1% | 94.6% | 94.9% |
All three remove refusals and keep MMLU within noise (our plain Q8_0 of the
original scores 86.70% and sits at KLD 0.017). Both of the others change the
model less than this release does, heretic-2 the least. What this release adds
is the form: an adapter for the quants you already run, and builds from 85 GB.
Everything was measured against one reference, one corpus, one machine: the
original BF16 model's own logits over a held-out neutral set, 87 chunks at 4096
context (BF16 PPL 4.047), llama.cpp 980aef8c, 8x RTX PRO 6000 (sm_120), CUDA 13.
Limitations
- The refusal count reads the opening of each reply; a soft refusal phrased as
an answer would be missed. - Thinking at effort
xhighis not production-ready on this build: 16 of 30
reasoning traces hit the 3072-token budget without closing, and a few repeat
themselves in a loop instead of finishing (the original refuses in a few
lines, so it never got there). Effortlowcloses cleanly; a coming revision
with a per-layer edit is expected to fix this. - Vision was not tested. The
mmprojfiles here are the original's projector,
unmodified. - Not run yet on a Mac, in LM Studio or in Ollama.
- The direction comes from English prompts at one token position. Russian
refusals went with it; other languages were not measured.
Responsible use
This model has a learned refusal direction removed, so it will answer requests
the original declines. Nothing about it makes the output safe, correct or
lawful, and any deployment needs its own access controls and policy
enforcement.
Credits
Direction estimation, adapter build, quantization and evaluation by
nik.bogatyrev. Base model Qwen3.8-Flash-Next by
Qwen, under the Qwen Community License 1.0. The
method follows Arditi et al., Refusal in Language Models Is Mediated by a
Single Direction (2024).


