← back to catalog · registered 2026-10-08 14:58

AtomicChat/Qwen3.8-Flash-Next-Abliterated-Uncensored-GGUF

AtomicChat Qwen GGUF MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/AtomicChat%2FQwen3.8-Flash-Next-Abliterated-Uncensored-GGUF"
Response includes
  • classification m8
  • files 8
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
1d ago
created 2026-10-07

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
gguf atomic-chat qwen qwen3.8 flash-next moe imatrix llama.cpp lora abliteration refusal-direction not-for-all-audiences

Related

Total size
0 B
Files
8
Quantizations
3
Registered
2026-10-08 14:58
Last updated on HF
2026-10-08 14:36

Files by quantization

BF16 1 file 866 MB
mmproj-Qwen3.8-Flash-Next-Abliterated-Uncensored-BF16.gguf 866 MB b115ede4 download
F16 1 file 862 MB
mmproj-Qwen3.8-Flash-Next-Abliterated-Uncensored-F16.gguf 862 MB 0e61454a download
Auxiliary files 6 files 138 KB
.gitattributes 68.3 KB 0b722cc9 download
btn_atomic.png 21.5 KB e4fc7a58 download
btn_discord.png 20.0 KB 17eac89b download
btn_github.png 14.1 KB 0482211b download
README.md 11.3 KB b4ab3d92 download
LICENSE 3.16 KB 9557a896 download

README current version from Hugging Face


license: other
license_name: qwen-community-1.0
license_link: https://huggingface.co/Qwen/Qwen3.8-Flash-Next/blob/main/LICENSE
base_model:

  • Qwen/Qwen3.8-Flash-Next
    base_model_relation: finetune
    quantized_by: AtomicChat
    pipeline_tag: text-generation
    library_name: gguf
    tags:
  • atomic-chat
  • qwen
  • qwen3.8
  • flash-next
  • moe
  • gguf
  • imatrix
  • llama.cpp
  • lora
  • abliteration
  • refusal-direction
  • not-for-all-audiences

How to Run Qwen3.8-Flash-Next Without Refusals Locally

One refusal direction projected out of Qwen's original weights, shipped two ways: three baked GGUF builds, and a 318 MB LoRA adapter for the Flash-Next quants you already have. The measurements behind this card are public.

Atomic Chat Discord GitHub
  • On held-out prompts, refusals fall from 92.6% to 1.2% in English and from 80% to 0% in Russian, and to none of 30 with thinking on. MMLU moves by 0.5 points, inside the noise.
  • The 93.9 GB build matches the original's token choice 91.1% of the time: closer to the original than our published 94.5 GB quant of the unmodified model (89.6%).
  • Already running our Flash-Next quants? Add the adapter with --lora-scaled instead of downloading a new model.

The builds

Build In memory On SSD Total Mean KLD Same top-1 Refused: EN / RU / XSTest
AD-3.86bpw-IQ4_XS 53.3 GB 32.0 GB 85.3 GB 0.113 87.8% 0% / 0% / 0.4%
AD-4.25bpw-Q4_K_M 61.9 GB 32.0 GB 93.9 GB 0.060 91.1% 0% / 0% / 0.8%
AD-4.85bpw-Q5_K_M 75.3 GB 32.0 GB 107.3 GB 0.037 93.0% 1.2% / 0% / 0.8%

AD-4.25bpw-Q4_K_M is the one to take if it fits. AD-3.86bpw is the one
whose in-memory part fits a 64 GB Mac next to the n-gram table on SSD, the same
way our 85 GB build
runs there; this build has not been run on a Mac yet.

KLD and top-1 are against the original, unmodified BF16 model, so they hold
the ablation and the quantization together. For scale: the ablation alone, on
BF16 weights, is 0.020; our plain Q8_0 of the original is 0.017.

Against the other way to get the same model, our original quant of the same
size with the adapter on top, the baked builds are closer to the original. The
gain is the layout, our current recipe, which the original quants predate; the
adapter itself adds about 0.007 to any quant:

Total Mean KLD Same top-1
baked, this repo: AD-3.86bpw-IQ4_XS 85.3 GB 0.113 87.8%
original quant AD-3.84bpw-IQ4_XS + adapter 84.9 GB 0.234 82.4%
baked, this repo: AD-4.25bpw-Q4_K_M 93.9 GB 0.060 91.1%
original quant AD-4.27bpw-Q4_K_M + adapter 94.5 GB 0.090 89.2%

Naming

Files are named by measured bits per weight, with the type tags of our
original quants, so
one tag means one size in both repos: IQ4_XS ~85 GB, Q4_K_M ~94 GB, Q5_K_M
~107-110 GB. The tag is not the experts' type: in AD-4.25bpw-Q4_K_M they are
IQ3_XXS and IQ4_XS, ffn_down_exps IQ4_NL, the n-gram table Q4_1.

What changed, and what did not

Original against ablated, both BF16, on prompts never used to pick anything:

Original Ablated
JBB harmful behaviours (EN, 81), refused 92.6% 1.2%
Aya red-teaming (RU, 100), refused 80.0% 0%
XSTest safe prompts (250), refused 5.2% 0.8%
Thinking low / xhigh (30 prompts), refused 97% / 80% 0% / 0%
MMLU, 2000 questions 87.30% 86.80% (-0.5, McNemar p = 0.31)
Tool calls (20) 20/20 valid 20/20 valid
Needle at 30k tokens (3 depths) 3/3 3/3
Mean KLD to the original, neutral / code - 0.020 / 0.017

Read by hand, the answers are complete and in the language asked. Refusal is
counted by the opening of the reply, so it is indicative, not a judge; empty or
degenerate replies count as damage, and there were none.

The adapter

lora/Qwen3.8-Flash-Next-Abliterated-Uncensored-LoRA.gguf: rank 1, 318 MB, 146 sites. It
applies to any GGUF of the original Flash-Next weights, including our three
published quants:

llama-server -m Qwen3.8-Flash-Next-AD-4.27bpw-Q4_K_M-M64-00001-of-00033.gguf \
  -ngl 99 --jinja -fit off \
  --lora-scaled lora/Qwen3.8-Flash-Next-Abliterated-Uncensored-LoRA.gguf:1.25

Or load it once and set the strength per request:

llama-server -m <model> -ngl 99 --jinja -fit off \
  --lora lora/Qwen3.8-Flash-Next-Abliterated-Uncensored-LoRA.gguf --lora-init-without-apply
curl localhost:8080/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "messages": [{"role": "user", "content": "..."}],
  "lora": [{"id": 0, "scale": 1.25}]
}'

Use scale 1.25. At 1.0, 47% of English validation prompts still refused;
above 1.25 the model moves further from the original with nothing left to
remove. On our published quants:

Published quant Mean KLD without / with Refused with: EN / RU / XSTest
AD-3.84bpw-IQ4_XS 0.227 / 0.234 1.2% / 0% / 0.4%
AD-4.27bpw-Q4_K_M 0.083 / 0.090 0% / 0% / 0.8%
AD-5.00bpw-Q5_K_M 0.083 / 0.089 0% / 0% / 0.8%

Get started

  • Atomic Chat:
    the easiest path. Open the app, search AtomicChat/Qwen3.8-Flash-Next-Abliterated-Uncensored-GGUF,
    pick a build, hit Use this model.
  • llama.cpp: download one build and point -m at its first shard.
hf download AtomicChat/Qwen3.8-Flash-Next-Abliterated-Uncensored-GGUF \
  --include "Qwen3.8-Flash-Next-Abliterated-Uncensored-AD-4.25bpw-Q4_K_M/*" --local-dir .
llama-server \
  -m Qwen3.8-Flash-Next-Abliterated-Uncensored-AD-4.25bpw-Q4_K_M/Qwen3.8-Flash-Next-Abliterated-Uncensored-AD-4.25bpw-Q4_K_M-00001-of-00036.gguf \
  -ngl 99 -c 32768 --jinja -fit off

As with the original quants:
keep mmap on, pass -fit off, and always --jinja for the model's own chat
template. Shard 2 of every build holds nothing but the n-gram table, so it can
stay on SSD. Needs a llama.cpp build with Qwen3.8-Flash-Next support
(PR #27742).

Sampling as Qwen recommends for the original: temperature 1.0, top_p 0.95,
top_k 20 with thinking; temperature 0.7, top_p 0.8, top_k 20,
presence_penalty 1.5 without.

How it was made

  1. Direction. Difference of means of the residual stream at the last prompt
    token, chat template applied, thinking off: 416 harmful against 416 harmless
    English prompts. Taken entering block 34 of 48, averaged over the four
    hyper-connection streams, with the harmless-mean component removed.
  2. Edit. W' = W - 1.25 r rᵀW on every matrix that writes into the
    residual: 48 attention / DeltaNet outputs, all 512 experts of all 48 MoE
    blocks, 48 shared experts, the n-gram value projection and the token
    embedding. The router and the hyper-connection weights are untouched.
  3. Choice. 34 variants screened on validation prompts (row, strength, which
    writers, one direction per block, English only or English + Russian); the
    numbers on this card come from held-out test prompts.
  4. Bake. The same edit applied to the BF16 weights in f32 and rounded once;
    every one of the 146 writers checked to keep 0.25 of its original component
    along r, within 3.4e-4. Then quantized with the importance matrix of our
    original quants.

One thing that did not work: editing the attention outputs alone, which is all
that Heretic-style tools reach on this architecture, left 96% of English
refusals in place with this direction. The refusal lives in the experts too.

Against other releases

Measured here at Q8_0, against the same BF16 reference and on the same prompts:

This release (Q8_0 + adapter) OrcaRouter Uncensored heretic-2
Refused, EN / RU 0% / 0% 0% / 1% 0% / 0%
XSTest refused 1.2% 0.8% 0.4%
MMLU (2000) 86.50% 86.65% 87.20%
Mean KLD to the original 0.026 0.022 0.019
Same top-1 94.1% 94.6% 94.9%

All three remove refusals and keep MMLU within noise (our plain Q8_0 of the
original scores 86.70% and sits at KLD 0.017). Both of the others change the
model less than this release does, heretic-2 the least. What this release adds
is the form: an adapter for the quants you already run, and builds from 85 GB.

Everything was measured against one reference, one corpus, one machine: the
original BF16 model's own logits over a held-out neutral set, 87 chunks at 4096
context (BF16 PPL 4.047), llama.cpp 980aef8c, 8x RTX PRO 6000 (sm_120), CUDA 13.

Limitations

  • The refusal count reads the opening of each reply; a soft refusal phrased as
    an answer would be missed.
  • Thinking at effort xhigh is not production-ready on this build: 16 of 30
    reasoning traces hit the 3072-token budget without closing, and a few repeat
    themselves in a loop instead of finishing (the original refuses in a few
    lines, so it never got there). Effort low closes cleanly; a coming revision
    with a per-layer edit is expected to fix this.
  • Vision was not tested. The mmproj files here are the original's projector,
    unmodified.
  • Not run yet on a Mac, in LM Studio or in Ollama.
  • The direction comes from English prompts at one token position. Russian
    refusals went with it; other languages were not measured.

Responsible use

This model has a learned refusal direction removed, so it will answer requests
the original declines. Nothing about it makes the output safe, correct or
lawful, and any deployment needs its own access controls and policy
enforcement.

Credits

Direction estimation, adapter build, quantization and evaluation by
nik.bogatyrev. Base model Qwen3.8-Flash-Next by
Qwen, under the Qwen Community License 1.0. The
method follows Arditi et al., Refusal in Language Models Is Mediated by a
Single Direction
(2024).

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration