license: apache-2.0
tags:
- abliteration
- heretic
- mixture-of-experts
- moe
- qwen3_5_moe
- method
Fused-MoE-expert abliteration for Heretic
A patch + method for abliterating fused-MoE models (e.g. Qwen3.6-35B-A3B / qwen3_5_moe) with Heretic, where the stock tool silently skips the experts and produces a weak, partial abliteration.
This is the method behind SparkyForge/Ember (BF16) and SparkyForge/Cinder (NVFP4): 5/100 refusals @ KL 0.0076 (94% reduction) with measured zero capability loss vs the base.
The problem
Heretic (and most abliteration tooling) finds a model's ablatable Linear layers by walking modules and wrapping the leaves it recognizes. qwen3_5_moe doesn't store its experts as a list of Linear modules — the 256 routed experts per layer are packed into one fused 3D nn.Parameter (down_proj shaped [num_experts, hidden, intermediate]) inside a Qwen3_5MoeSparseMoeBlock, alongside a dense shared_expert that runs on every token.
Result: the standard targeting finds only the attention o_proj for these layers. The entire MoE/MLP block is left un-abliterated. You get a model that looks abliterated (the search runs, KL moves) but still refuses, because the experts — where most of the FFN computation lives — were never touched. (This is why naive abliterations of this model land weak, ~60/100 refusals.)
There's a second, quieter failure: even if you add the fused block as a target, Heretic's abliterate() loop reaches into module.weight, and the fused block has no .weight → AttributeError: 'Qwen3_5MoeSparseMoeBlock' object has no attribute 'weight'.
The method
The patch (heretic-fused-experts.patch, included here) does three things:
Surfaces the fused MoE block as one ablitable component (
mlp.down_proj(fused)), excludes it from the LoRA/PEFT target set (it can't wrap a 3D Parameter), and abliterates both the routed experts and the denseshared_expert.Memory-safe reset via forward hooks. Abliterating
down_projbyW -= λ·v(vᵀW)is mathematically identical to a rank-1 projection of the MoE block's output:y -= λ·v(vᵀy). So instead of editing (and backing up) the 32GB of 3D expert weights for every trial, a single forward hook per layer reproduces routed + shared expert ablation exactly, for any strength λ — at ~0.7 MB of state. Reset = remove the hook. This is what makes Heretic's per-trial Optuna search over a 256-expert model tractable without OOM. The chosen direction/strength is baked into the weights once, at save time.Respects the hybrid layers. The 30 linear-attention (Mamba/GDN) layers are left untouched; only attention
o_proj+ the MoE block are abliterated.
It also adds the if component == "mlp.down_proj(fused)": continue guard in abliterate() so the fused block is handled solely by the hooks (fixing the AttributeError).
Apply
The patch targets Heretic v1.3.0 (heretic/{model.py,main.py}):
# inside your Heretic install/container, from the package root:
patch -p1 < heretic-fused-experts.patch
# then run Heretic normally on a qwen3_5_moe model
Then run the standard Heretic flow. Watch the log for Abliterable components: ... mlp.down_proj(fused): N modules — if you only see attn.o_proj, the patch didn't take.
Results
On Qwen3.6-35B-A3B: refusals 86 → 5 / 100 at KL 0.0076 to the base; a 30-probe retention suite matched the base on every dimension across N=10 runs (extraction, multi-hop, reasoning, arithmetic, factual, code, language, instruction, format). Weights: Ember (BF16), Cinder (NVFP4).
Attribution & license
- Method + patch: SparkyForge.
- Built on: Heretic by Philipp Emanuel Weidmann — apply on top of Heretic v1.3.0; the patched portions are subject to Heretic's upstream license (see its repo).
- Model: Qwen/Qwen3.6-35B-A3B (Apache 2.0), © the Qwen team.
Independent community work; not affiliated with NVIDIA, the Apache Software Foundation, the Qwen team, or the Heretic project.