license: other
license_name: qwen-community-license-1.0
license_link: https://huggingface.co/haihengh/Qwen3.8-Flash-Next-125B-finchmoe-4bit-ple4bit-abliterated/blob/main/LICENSE
base_model: Qwen/Qwen3.8-Flash-Next
language: [en]
tags: [qwen, moe, abliterated, uncensored, apple-silicon, metal, quantization, finchmoe, ssd-streaming]
library_name: finchmoe
Qwen3.8-Flash-Next-125B (abliterated) — FinchMoE 4-bit, quantized PLE
An abliterated (refusal-direction-removed) build of Qwen3.8-Flash-Next-125B
— a 48-layer MoE with 512 experts, top-8 — repacked by
FinchMoE into the .finch format for
SSD-streaming inference on memory-constrained Apple Silicon. Every tensor class
is int4/int8 affine, including the PLE n-gram table. 96.9 GiB on disk, down
from ~360 GB in BF16.
This model is abliterated
The refusal direction has been removed from the weights by the upstream
abliteration, not by prompting or by any change to the runtime. Expect
substantially fewer refusals than the base model.
Refusal behaviour itself is not measured here. What is measured is that the
process did not damage general capability — see below. Nothing in this repository
was evaluated for what it will or will not refuse.
The upstream release ships its abliteration scripts (apply_ablation_flashnext.py,capture_refusal_flashnext.py, verify_ablit_flashnext.py) and ABLIT_META.json
alongside the weights. Those are the authority on which tensors were modified and
how the direction was chosen; this repository only repacks the result.
Measured quality
EvalPlus HumanEval, greedy, 164 problems, the project's frozen server protocol,
measured 2026-09-27. The two rows differ only in the weights: same engine,
same harness, same quantization settings, same context (4096).
| base pass@1 | HumanEval+ | |
|---|---|---|
| Qwen3.8-Flash-Next-125B (base weights) | 0.9451 (155/164) | 0.9207 (151/164) |
| this abliterated build | 0.9451 (155/164) | 0.9268 (152/164) |
Identical on the base suite and one problem better on the stricter suite — well
inside noise. Abliteration did not measurably degrade coding ability.
The residual failures are dominated by the harness, not the model: of the 9
failing problems, 7 are length-capped by the 768-token budget
(HumanEval/32, 93, 113, 116, 129, 130, 132) and only 2 are genuine (145, 163).
So this build's real capability failure set on HumanEval is 2 problems, and both
of those fail on the base weights as well — the same stable core that fails
across every Qwen install in this project.
Provenance
| upstream abliteration | windowsxp811203/Qwen3.8-Flash-Next-Abliterated |
| base weights | Qwen/Qwen3.8-Flash-Next |
| repacked by | FinchMoE FinchMoERepack |
The modifications from the abliterated snapshot are the repack into .finch and
the quantization described below — nothing else. The upstream snapshot was taken
as-is; this repository does not re-abliterate or fine-tune.
Two hashes identify this install, and the distinction matters if you are
comparing it against the base release:
sourceSnapshotHash sha256:99e81524…590de— identical to the base release's.
That hash covers the tensor index (names, shapes, layout), which abliteration
does not change. Two installs sharing it are structurally identical, not
materially identical.model_weights.binsha2566af82b557d8207e470f50c1fba5dbb14e7ff4b48139a5b0b2e4f8ac56d188acb
— this is what actually differs from the base release'sc522877f…2a06.
Quantization
| tensor class | bits | group |
|---|---|---|
| routed experts, shared expert, attention, embeddings | 4 | 64 |
| linear-attention projections, router | 8 | 64 |
| PLE n-gram table | 4 | 32 |
Each PLE row is laid out fixed-stride as [packed nibbles][BF16 scales][BF16 biases], so one pread fetches a row's whole quantized state and the decoder
derives the group size from the row itself.
Files
| File | Size | Content |
|---|---|---|
model_weights.bin |
3.9 GB | non-expert weights (attention, hyper-connections, norms, embeddings) |
packed_experts/ |
68.1 GB | routed + shared experts, 48 layers, 512 experts each |
ple_shards/ |
32.0 GB | PLE n-gram table, 128 shards, int4 affine group 32 |
manifest.json |
31 KB | tensor map, quantization slots, architecture |
tokenizer/ |
23 MB | tokenizer and chat template |
verified-install.json |
29 KB | install receipt (per-file SHA-256, lets the engine skip re-hashing) |
This is a FinchMoE-specific format, not GGUF, MLX or safetensors. It will not
load in llama.cpp, MLX or transformers.
How to run
Build FinchMoE and point it at this
directory:
FinchMoECLI --model /path/to/Qwen3.8-Flash-Next-125B-finchmoe-4bit-ple4bit-abliterated \
--prompt "..." --temperature 0.7
Or as an OpenAI-compatible server:
FinchMoEServer --model /path/to/Qwen3.8-Flash-Next-125B-finchmoe-4bit-ple4bit-abliterated \
--port 8080 --verify trusted-install
Expect roughly 19 tok/s prompt processing and ~2.5 tok/s generation on a 16 GB
M4 Mac mini with the install streaming from an external SSD. Resident memory is
dominated by the KV cache and scratch buffers, not by the weights — the expert
and PLE tensors are streamed per token.
Note --verify trusted-install skips the first-touch SHA-256 over the whole
install (which is ~167 GB of files and takes a minute on first load); use--verify strict if you want every byte checked.
License
The base model is licensed under the Qwen Community License 1.0, which is
included here as LICENSE and permits publishing, distributing and creating
derivative works subject to its conditions. This repository is a derivative work:
the modifications are the repack into the .finch format and the quantization
described above. The copyright and permission notice is retained in full.