← back to catalog · registered 2026-09-27 23:57

haihengh/Qwen3.6-35B-A3B-finchmoe-4bit-abliterated

haihengh Qwen 35B MoE
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/haihengh%2FQwen3.6-35B-A3B-finchmoe-4bit-abliterated"
Response includes
  • classification m-uncensored
  • files 2
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-09-27
Downloads over time
Now0→from0↑0%
00110 on Sep 270 on Sep 28Sep
Sep 27 → Sep 28 · 2 snapshots · spans 1 day

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
finchmoe qwen moe abliterated uncensored apple-silicon metal quantization en base_model:Qwen/Qwen3.6-35B-A3B base_model:finetune:Qwen/Qwen3.6-35B-A3B license:apache-2.0

Related

Total size
0 B
Files
2
Quantizations
1
Registered
2026-09-27 23:57
Last updated on HF
2026-09-27 23:20

Files by quantization

Auxiliary files 2 files 6.75 KB
README.md 5.27 KB eaa4eaf4 download
.gitattributes 1.48 KB a6344aac download

README current version from Hugging Face


license: apache-2.0
base_model: Qwen/Qwen3.6-35B-A3B
language: [en]
tags: [qwen, moe, abliterated, uncensored, apple-silicon, metal, quantization, finchmoe]
library_name: finchmoe

Qwen3.6-35B-A3B (abliterated) — FinchMoE 4-bit (Apple Silicon)

An abliterated (refusal-direction-removed) build of Qwen3.6-35B-A3B, repacked
by FinchMoE into the .finch format for
SSD-streaming inference on memory-constrained Apple Silicon. 18.7 GiB on disk,
down from ~72 GB in BF16.

This model is abliterated

The refusal direction has been removed from the weights by the upstream
abliteration, not by prompting or by any change to the runtime. Expect
substantially fewer refusals than the base model.

Refusal behaviour itself is not measured here. What is measured is that the
process did not damage general capability — see below. Nothing in this repository
was evaluated for what it will or will not refuse.

Measured quality

EvalPlus HumanEval, greedy, 164 problems, the project's frozen server protocol,
measured 2026-09-27. The two rows differ only in the weights: same engine,
same harness, same quantization settings, same context (4096).

base pass@1 HumanEval+
Qwen3.6-35B-A3B (base weights) 0.9085 (149/164) 0.8780 (144/164)
this abliterated build 0.9329 (153/164) 0.9024 (148/164)

Both suites improve by 4 problems. Do not read that as abliteration improving
coding ability.
The binomial standard error at p≈0.91, n=164 is 2.2 points, so
+2.4 pt is about one standard error — it is noise. Problems were gained and
lost in both directions (gained 95, 99, 113, 124, 147, 160; lost 54, 116). The
defensible claim is no measurable regression, which is the thing worth
establishing for an abliterated derivative.

One caveat on the residual failures: of the 11 problems that fail, 4 are
length-capped by the harness's 768-token budget (HumanEval/116, 129, 130, 132)
rather than genuinely wrong, and 7 fail on their own merits (32, 54, 62, 93, 134,
145, 163). A stable core — 32, 93, 129, 130, 132, 145, 163 — fails on the base
weights too, so it is model difficulty, not abliteration.

Provenance

upstream abliteration huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated
base weights Qwen/Qwen3.6-35B-A3B
repacked by FinchMoE FinchMoERepack

The modifications from the abliterated snapshot are the repack into .finch and
the quantization described below — nothing else. The upstream snapshot was taken
as-is; this repository does not re-abliterate or fine-tune.

Two hashes identify this install, and the distinction matters if you are
comparing it against the base release:

  • sourceSnapshotHash sha256:41b93561…0be83 — identical to the base release's.
    That hash covers the tensor index (names, shapes, layout), which abliteration
    does not change. Two installs sharing it are structurally identical, not
    materially identical.
  • model_weights.bin sha256 f6862341c9688e234c682cef186af5a92445be1dbcdd36d59637338634d311bd
    — this is what actually differs from the base release's 9644b61a…8228d.

Quantization

Every tensor class is affine with BF16 scales and biases, group size 64 —
byte-for-byte the same scheme as the base 4-bit release:

tensor class bits
routed experts, shared expert, attention, embeddings 4
linear-attention (GDN) projections, router 8
norms, gates BF16

Files

File Size Content
model_weights.bin 1.9 GB non-expert weights (attention, GDN, embeddings, norms)
packed_experts/ 18.1 GB routed + shared experts, 40 layers, 256 experts each
manifest.json 10 KB tensor map, quantization slots, architecture
tokenizer/ 23 MB tokenizer and chat template
verified-install.json 8 KB install receipt (per-file SHA-256, lets the engine skip re-hashing)

This is a FinchMoE-specific format, not GGUF, MLX or safetensors. It will not
load in llama.cpp, MLX or transformers.

How to run

Build FinchMoE and point it at this
directory:

FinchMoECLI --model /path/to/Qwen3.6-35B-A3B-finchmoe-4bit-abliterated \
            --prompt "..." --temperature 0.7

Or as an OpenAI-compatible server:

FinchMoEServer --model /path/to/Qwen3.6-35B-A3B-finchmoe-4bit-abliterated \
               --port 8080 --verify trusted-install

Expect roughly 40 tok/s prompt processing and ~8 tok/s generation on a 16 GB M4
Mac mini with the install streaming from an external SSD. Resident memory is
dominated by the KV cache and scratch buffers, not by the weights — the expert
tensors are streamed per token.

Note --verify trusted-install skips the first-touch SHA-256 over the whole
install; use --verify strict if you want every byte checked on load.

License

Apache-2.0, inherited from Qwen3.6-35B-A3B. This repository is a derivative work:
the modifications are the repack into the .finch format and the quantization
above. The base model's licence and the upstream abliteration's terms carry
through unchanged.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.