← back to catalog · registered 2026-10-10 17:58

satellitedown/Qwen3.8-Flash-Next-abliterated-fngine

satellitedown Qwen MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/satellitedown%2FQwen3.8-Flash-Next-abliterated-fngine"
Response includes
  • classification m-uncensored
  • files 4
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-10

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
abliterated uncensored qwen3.8 flash-next mixture-of-experts gsq rco fp8 cuda fngine text-generation arxiv:2604.18556

Related

Total size
0 B
Files
4
Quantizations
1
Registered
2026-10-10 17:58
Last updated on HF
2026-10-10 17:43

Files by quantization

Auxiliary files 4 files 13.1 KB
README.md 7.56 KB 847202ad download
LICENSE 3.16 KB 9557a896 download
.gitattributes 1.55 KB d7d3d896 download
NOTICE 857 B faf23ad4 download

README current version from Hugging Face


license: other
license_name: qwen-community-1.0
license_link: LICENSE
base_model:

  • Qwen/Qwen3.8-Flash-Next
    pipeline_tag: text-generation
    tags:
  • abliterated
  • uncensored
  • qwen3.8
  • flash-next
  • mixture-of-experts
  • gsq
  • rco
  • fp8
  • cuda
  • fngine

Qwen3.8-Flash-Next · abliterated · fngine pack

An abliterated build of Qwen/Qwen3.8-Flash-Next in the
format of fngine, a single-GPU C++20/CUDA engine written for
this model (source at commit 9adf5d6b3647 included in fngine/).
It is the same model as the
ISTA-DASLab GSQ-RCO IQ3_S
quantization, carrying the abliteration of
huihui-ai/Huihui-Qwen3.8-Flash-Next-abliterated.

Reduced safety filtering. Refusal behaviour has been removed. The model can produce content
that the original would decline. You are responsible for how you use it and for complying with
the law and the license.

What is in this repository

Path Contents Size
pack/ the fngine pack: FP8 dense weights, GSQ-RCO routed experts (native GGUF i-quant bytes), IQ4_NL n-gram table, IQ4_NL MTP experts, tokenizer, placement profiles, ablation.bin ~89 GiB
fngine/ engine source at commit 9adf5d6b3647 (CMake, CUDA 13.3, sm_120a), tests and tooling ~9 MB
abliteration/ how the abliteration was measured and transferred, with all verification evidence small

pack/dense_bf16.bin (8.1 GiB) is only used by fngine's validation mode; serving does not need it.

How the abliteration was transferred

The released abliteration was measured as a weight delta against the official release (both
range-fetched from the Hub):

  • Only the residual-stream writers changed: in each of the 48 layers, the attention/GDN output
    projection, the shared-expert down projection, and every routed expert's down projection.
    Embeddings, lm_head, router, hyper-connections, n-gram/PLE, Q/K/V, indexer and the MTP head are
    byte-identical.
  • Every changed tensor satisfies W' = W − α·r·rᵀW with a single refusal direction r
    (minimum cosine 0.9999985 across the 48 per-layer fits) and α ≈ 1.955 (the remaining ≤ 4.6% of
    each delta is BF16 rounding). With α ≈ 2 the refusal component is reflected, not merely removed.

In this pack:

  • the 96 changed dense tensors are the abliterated release's exact BF16 bytes, converted to
    fngine's FP8 rows by the same converter as every other tensor (804 other dense tensors are
    byte-identical to the un-abliterated pack);
  • the routed experts keep their GSQ-RCO bytes. The routed output is a weighted sum of linear down
    projections, so editing every expert is exactly equivalent to editing their sum:
    y_routed −= α·r·(r·y_routed). fngine applies this in its MoE combine kernels (decode, verify and
    prefill; GPU- and CPU-resident experts) from pack/ablation.bin. The shared expert is added
    unedited because its weights already carry the edit;
  • everything else is unchanged from the un-abliterated pack.

fngine-serve --no-ablation ignores ablation.bin (dense edit only), for comparison.

Verification

Measured on an RTX 5090 against the un-abliterated pack; details and raw data in
abliteration/README.md and abliteration/results/.

Agreement with independent PyTorch references (official HF modeling code, the same GGUF experts):

Run G1 (2,048 tokens) top-1 / KLD G2 (8,192 tokens) top-1 / KLD
Un-abliterated pack vs original reference 0.815–0.831 / 0.097–0.122 0.908–0.909 / 0.030–0.032
This pack vs abliterated reference 0.873–0.896 / 0.052–0.067 0.914–0.918 / 0.026–0.028
This pack, routed-expert edit off 0.729 / 0.345 0.806 / 0.188

The gap to 1.0 is the FP8 dense weights and 8-bit KV cache, and is the same for both packs.

Refusals on 32 mild prompts that safety-tuned models often decline (creative writing, profanity,
satire, security-training examples): un-abliterated 8 refused, this pack 0, with none newly refused.

Capability: math 20/20 (exact answers), Python 12/12 (executed against unit tests) and knowledge
12/12 for both packs. Speed: no measurable difference (five paired workloads, all within one
standard deviation). Engine tests: 46/46 pass on this pack.

Performance (RTX 5090, single user)

Decode, mean of 6 repetitions, speculative decoding on, 256K context configured:

Workload Decode tok/s
Coding chat (512 generated tokens), turn 1 / turn 2 245 / 269
Document continuation after a 1,024-token prompt 201
Prefill, 4,096-token prompt ~2,860 tok/s

Absolute numbers depend on GPU clock, host DRAM bandwidth, PCIe width and other GPU load.

Requirements

  • NVIDIA RTX 5090 (or another sm_120a GPU with 32 GB); fngine is compiled for that architecture only
  • an x86-64 CPU with AVX-512 VNNI, BF16 and VBMI (AMD Zen 4/5, Intel Sapphire Rapids or later) for the
    CPU expert tier
  • ~60 GB host RAM (CPU-tier experts, embedding, n-gram cache) and an NVMe SSD (the n-gram table is read on demand)
  • Linux x86_64 with a working NVIDIA driver; ~87 GB of disk for the serving files

Usage

The installer builds fngine at a
pinned commit with an isolated CUDA toolchain, downloads and verifies this pack (without the
validation-only dense_bf16.bin), starts the server and updates both:

git clone https://github.com/satellitedown/flash-next-abliterated-fngine
cd flash-next-abliterated-fngine
bash setup.sh

Manually (CUDA 13.3, CMake ≥ 3.28, Ninja, liburing, a C++20 compiler and a CUDA-supported host compiler):

hf download satellitedown/Qwen3.8-Flash-Next-abliterated-fngine --local-dir qwen3.8-flash-next-abliterated \
    --exclude pack/dense_bf16.bin
cd qwen3.8-flash-next-abliterated/fngine
cmake --preset release && cmake --build build -j
build/apps/fngine-serve --pack ../pack --ctx 262144 --host 127.0.0.1 --port 8002 \
    --model-id qwen3.8-flash-next

OpenAI-compatible Chat Completions (streaming, tools, reasoning_content, enable_thinking):

curl -N http://127.0.0.1:8002/v1/chat/completions -H 'Content-Type: application/json' \
  -d '{"model":"qwen3.8-flash-next","stream":true,"messages":[{"role":"user","content":"Hello"}]}'

The server runs one request at a time with a short FIFO queue. With agent frameworks that issue
parallel requests, cap concurrency to 1 for this endpoint. See fngine/README.md for every option,
including 256K-context notes, speculative decoding and validation tooling.

Credits and licenses

  • Qwen3.8-Flash-Next: Qwen, Qwen Community License 1.0. This repository is a derivative
    work; the license and its conditions apply to it.
  • GSQ-RCO quantization of the routed experts and n-gram table: ISTA-DASLab
    (GSQ, RCO); Apache-2.0,
    inheriting the base model's license.
  • Abliteration: huihui-ai
    (method: remove-refusals-with-transformers).
  • fngine: Apache-2.0; derived in part from Cinference and NInfer, with ggml i-quant code (MIT) and
    vendored libraries under their own licenses. See fngine/NOTICE and fngine/LICENSE.

Provided as is, without warranty. Do not expose the server unauthenticated to untrusted networks.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration