license: other
license_name: qwen-community-1.0
license_link: LICENSE
base_model:
- Qwen/Qwen3.8-Flash-Next
pipeline_tag: text-generation
tags: - abliterated
- uncensored
- qwen3.8
- flash-next
- mixture-of-experts
- gsq
- rco
- fp8
- cuda
- fngine
Qwen3.8-Flash-Next · abliterated · fngine pack
An abliterated build of Qwen/Qwen3.8-Flash-Next in the
format of fngine, a single-GPU C++20/CUDA engine written for
this model (source at commit 9adf5d6b3647 included in fngine/).
It is the same model as the
ISTA-DASLab GSQ-RCO IQ3_S
quantization, carrying the abliteration of
huihui-ai/Huihui-Qwen3.8-Flash-Next-abliterated.
Reduced safety filtering. Refusal behaviour has been removed. The model can produce content
that the original would decline. You are responsible for how you use it and for complying with
the law and the license.
What is in this repository
| Path | Contents | Size |
|---|---|---|
pack/ |
the fngine pack: FP8 dense weights, GSQ-RCO routed experts (native GGUF i-quant bytes), IQ4_NL n-gram table, IQ4_NL MTP experts, tokenizer, placement profiles, ablation.bin |
~89 GiB |
fngine/ |
engine source at commit 9adf5d6b3647 (CMake, CUDA 13.3, sm_120a), tests and tooling |
~9 MB |
abliteration/ |
how the abliteration was measured and transferred, with all verification evidence | small |
pack/dense_bf16.bin (8.1 GiB) is only used by fngine's validation mode; serving does not need it.
How the abliteration was transferred
The released abliteration was measured as a weight delta against the official release (both
range-fetched from the Hub):
- Only the residual-stream writers changed: in each of the 48 layers, the attention/GDN output
projection, the shared-expert down projection, and every routed expert's down projection.
Embeddings,lm_head, router, hyper-connections, n-gram/PLE, Q/K/V, indexer and the MTP head are
byte-identical. - Every changed tensor satisfies
W' = W − α·r·rᵀWwith a single refusal directionr
(minimum cosine 0.9999985 across the 48 per-layer fits) and α ≈ 1.955 (the remaining ≤ 4.6% of
each delta is BF16 rounding). With α ≈ 2 the refusal component is reflected, not merely removed.
In this pack:
- the 96 changed dense tensors are the abliterated release's exact BF16 bytes, converted to
fngine's FP8 rows by the same converter as every other tensor (804 other dense tensors are
byte-identical to the un-abliterated pack); - the routed experts keep their GSQ-RCO bytes. The routed output is a weighted sum of linear down
projections, so editing every expert is exactly equivalent to editing their sum:y_routed −= α·r·(r·y_routed). fngine applies this in its MoE combine kernels (decode, verify and
prefill; GPU- and CPU-resident experts) frompack/ablation.bin. The shared expert is added
unedited because its weights already carry the edit; - everything else is unchanged from the un-abliterated pack.
fngine-serve --no-ablation ignores ablation.bin (dense edit only), for comparison.
Verification
Measured on an RTX 5090 against the un-abliterated pack; details and raw data inabliteration/README.md and abliteration/results/.
Agreement with independent PyTorch references (official HF modeling code, the same GGUF experts):
| Run | G1 (2,048 tokens) top-1 / KLD | G2 (8,192 tokens) top-1 / KLD |
|---|---|---|
| Un-abliterated pack vs original reference | 0.815–0.831 / 0.097–0.122 | 0.908–0.909 / 0.030–0.032 |
| This pack vs abliterated reference | 0.873–0.896 / 0.052–0.067 | 0.914–0.918 / 0.026–0.028 |
| This pack, routed-expert edit off | 0.729 / 0.345 | 0.806 / 0.188 |
The gap to 1.0 is the FP8 dense weights and 8-bit KV cache, and is the same for both packs.
Refusals on 32 mild prompts that safety-tuned models often decline (creative writing, profanity,
satire, security-training examples): un-abliterated 8 refused, this pack 0, with none newly refused.
Capability: math 20/20 (exact answers), Python 12/12 (executed against unit tests) and knowledge
12/12 for both packs. Speed: no measurable difference (five paired workloads, all within one
standard deviation). Engine tests: 46/46 pass on this pack.
Performance (RTX 5090, single user)
Decode, mean of 6 repetitions, speculative decoding on, 256K context configured:
| Workload | Decode tok/s |
|---|---|
| Coding chat (512 generated tokens), turn 1 / turn 2 | 245 / 269 |
| Document continuation after a 1,024-token prompt | 201 |
| Prefill, 4,096-token prompt | ~2,860 tok/s |
Absolute numbers depend on GPU clock, host DRAM bandwidth, PCIe width and other GPU load.
Requirements
- NVIDIA RTX 5090 (or another
sm_120aGPU with 32 GB); fngine is compiled for that architecture only - an x86-64 CPU with AVX-512 VNNI, BF16 and VBMI (AMD Zen 4/5, Intel Sapphire Rapids or later) for the
CPU expert tier - ~60 GB host RAM (CPU-tier experts, embedding, n-gram cache) and an NVMe SSD (the n-gram table is read on demand)
- Linux x86_64 with a working NVIDIA driver; ~87 GB of disk for the serving files
Usage
The installer builds fngine at a
pinned commit with an isolated CUDA toolchain, downloads and verifies this pack (without the
validation-only dense_bf16.bin), starts the server and updates both:
git clone https://github.com/satellitedown/flash-next-abliterated-fngine
cd flash-next-abliterated-fngine
bash setup.sh
Manually (CUDA 13.3, CMake ≥ 3.28, Ninja, liburing, a C++20 compiler and a CUDA-supported host compiler):
hf download satellitedown/Qwen3.8-Flash-Next-abliterated-fngine --local-dir qwen3.8-flash-next-abliterated \
--exclude pack/dense_bf16.bin
cd qwen3.8-flash-next-abliterated/fngine
cmake --preset release && cmake --build build -j
build/apps/fngine-serve --pack ../pack --ctx 262144 --host 127.0.0.1 --port 8002 \
--model-id qwen3.8-flash-next
OpenAI-compatible Chat Completions (streaming, tools, reasoning_content, enable_thinking):
curl -N http://127.0.0.1:8002/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"model":"qwen3.8-flash-next","stream":true,"messages":[{"role":"user","content":"Hello"}]}'
The server runs one request at a time with a short FIFO queue. With agent frameworks that issue
parallel requests, cap concurrency to 1 for this endpoint. See fngine/README.md for every option,
including 256K-context notes, speculative decoding and validation tooling.
Credits and licenses
- Qwen3.8-Flash-Next: Qwen, Qwen Community License 1.0. This repository is a derivative
work; the license and its conditions apply to it. - GSQ-RCO quantization of the routed experts and n-gram table: ISTA-DASLab
(GSQ, RCO); Apache-2.0,
inheriting the base model's license. - Abliteration: huihui-ai
(method: remove-refusals-with-transformers). - fngine: Apache-2.0; derived in part from Cinference and NInfer, with ggml i-quant code (MIT) and
vendored libraries under their own licenses. Seefngine/NOTICEandfngine/LICENSE.
Provided as is, without warranty. Do not expose the server unauthenticated to untrusted networks.