license: apache-2.0
base_model: Qwen/Qwen3.6-35B-A3B
language: [en]
tags: [qwen, moe, abliterated, uncensored, apple-silicon, metal, quantization, finchmoe]
library_name: finchmoe
Qwen3.6-35B-A3B (abliterated) — FinchMoE 4-bit (Apple Silicon)
An abliterated (refusal-direction-removed) build of Qwen3.6-35B-A3B, repacked
by FinchMoE into the .finch format for
SSD-streaming inference on memory-constrained Apple Silicon. 18.7 GiB on disk,
down from ~72 GB in BF16.
This model is abliterated
The refusal direction has been removed from the weights by the upstream
abliteration, not by prompting or by any change to the runtime. Expect
substantially fewer refusals than the base model.
Refusal behaviour itself is not measured here. What is measured is that the
process did not damage general capability — see below. Nothing in this repository
was evaluated for what it will or will not refuse.
Measured quality
EvalPlus HumanEval, greedy, 164 problems, the project's frozen server protocol,
measured 2026-09-27. The two rows differ only in the weights: same engine,
same harness, same quantization settings, same context (4096).
| base pass@1 | HumanEval+ | |
|---|---|---|
| Qwen3.6-35B-A3B (base weights) | 0.9085 (149/164) | 0.8780 (144/164) |
| this abliterated build | 0.9329 (153/164) | 0.9024 (148/164) |
Both suites improve by 4 problems. Do not read that as abliteration improving
coding ability. The binomial standard error at p≈0.91, n=164 is 2.2 points, so
+2.4 pt is about one standard error — it is noise. Problems were gained and
lost in both directions (gained 95, 99, 113, 124, 147, 160; lost 54, 116). The
defensible claim is no measurable regression, which is the thing worth
establishing for an abliterated derivative.
One caveat on the residual failures: of the 11 problems that fail, 4 arelength-capped by the harness's 768-token budget (HumanEval/116, 129, 130, 132)
rather than genuinely wrong, and 7 fail on their own merits (32, 54, 62, 93, 134,
145, 163). A stable core — 32, 93, 129, 130, 132, 145, 163 — fails on the base
weights too, so it is model difficulty, not abliteration.
Provenance
| upstream abliteration | huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated |
| base weights | Qwen/Qwen3.6-35B-A3B |
| repacked by | FinchMoE FinchMoERepack |
The modifications from the abliterated snapshot are the repack into .finch and
the quantization described below — nothing else. The upstream snapshot was taken
as-is; this repository does not re-abliterate or fine-tune.
Two hashes identify this install, and the distinction matters if you are
comparing it against the base release:
sourceSnapshotHash sha256:41b93561…0be83— identical to the base release's.
That hash covers the tensor index (names, shapes, layout), which abliteration
does not change. Two installs sharing it are structurally identical, not
materially identical.model_weights.binsha256f6862341c9688e234c682cef186af5a92445be1dbcdd36d59637338634d311bd
— this is what actually differs from the base release's9644b61a…8228d.
Quantization
Every tensor class is affine with BF16 scales and biases, group size 64 —
byte-for-byte the same scheme as the base 4-bit release:
| tensor class | bits |
|---|---|
| routed experts, shared expert, attention, embeddings | 4 |
| linear-attention (GDN) projections, router | 8 |
| norms, gates | BF16 |
Files
| File | Size | Content |
|---|---|---|
model_weights.bin |
1.9 GB | non-expert weights (attention, GDN, embeddings, norms) |
packed_experts/ |
18.1 GB | routed + shared experts, 40 layers, 256 experts each |
manifest.json |
10 KB | tensor map, quantization slots, architecture |
tokenizer/ |
23 MB | tokenizer and chat template |
verified-install.json |
8 KB | install receipt (per-file SHA-256, lets the engine skip re-hashing) |
This is a FinchMoE-specific format, not GGUF, MLX or safetensors. It will not
load in llama.cpp, MLX or transformers.
How to run
Build FinchMoE and point it at this
directory:
FinchMoECLI --model /path/to/Qwen3.6-35B-A3B-finchmoe-4bit-abliterated \
--prompt "..." --temperature 0.7
Or as an OpenAI-compatible server:
FinchMoEServer --model /path/to/Qwen3.6-35B-A3B-finchmoe-4bit-abliterated \
--port 8080 --verify trusted-install
Expect roughly 40 tok/s prompt processing and ~8 tok/s generation on a 16 GB M4
Mac mini with the install streaming from an external SSD. Resident memory is
dominated by the KV cache and scratch buffers, not by the weights — the expert
tensors are streamed per token.
Note --verify trusted-install skips the first-touch SHA-256 over the whole
install; use --verify strict if you want every byte checked on load.
License
Apache-2.0, inherited from Qwen3.6-35B-A3B. This repository is a derivative work:
the modifications are the repack into the .finch format and the quantization
above. The base model's licence and the upstream abliteration's terms carry
through unchanged.