← back to catalog · registered 2026-08-22 13:56

SparkyForge/heretic-fused-moe-abliteration

Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/SparkyForge%2Fheretic-fused-moe-abliteration"
Response includes
  • classification m3
  • files 4
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · 30-day
0
Likes
0
Model age
3mo ago
created 2026-06-23
Downloads over time
Now0→from0↑0%
00110 on Jun 240 on Oct 11JunJulAugSepOct
Jun 24 → Oct 11 · 55 snapshots · spans 109 days

Metadata

License
apache-2.0
Tags
abliteration heretic mixture-of-experts moe qwen3_5_moe method license:apache-2.0 region:us
Total size
0 B
Files
4
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-23 02:10

Files by quantization

Auxiliary files 4 files 30.9 KB
heretic-fused-experts.patch 24.4 KB 928ff21c download
README.md 4.35 KB 22f0ba1f download
.gitattributes 1.48 KB a6344aac download
NOTICE.txt 715 B 0414868c download

README current version from Hugging Face


license: apache-2.0
tags:

  • abliteration
  • heretic
  • mixture-of-experts
  • moe
  • qwen3_5_moe
  • method

Fused-MoE-expert abliteration for Heretic

A patch + method for abliterating fused-MoE models (e.g. Qwen3.6-35B-A3B / qwen3_5_moe) with Heretic, where the stock tool silently skips the experts and produces a weak, partial abliteration.

This is the method behind SparkyForge/Ember (BF16) and SparkyForge/Cinder (NVFP4): 5/100 refusals @ KL 0.0076 (94% reduction) with measured zero capability loss vs the base.

The problem

Heretic (and most abliteration tooling) finds a model's ablatable Linear layers by walking modules and wrapping the leaves it recognizes. qwen3_5_moe doesn't store its experts as a list of Linear modules — the 256 routed experts per layer are packed into one fused 3D nn.Parameter (down_proj shaped [num_experts, hidden, intermediate]) inside a Qwen3_5MoeSparseMoeBlock, alongside a dense shared_expert that runs on every token.

Result: the standard targeting finds only the attention o_proj for these layers. The entire MoE/MLP block is left un-abliterated. You get a model that looks abliterated (the search runs, KL moves) but still refuses, because the experts — where most of the FFN computation lives — were never touched. (This is why naive abliterations of this model land weak, ~60/100 refusals.)

There's a second, quieter failure: even if you add the fused block as a target, Heretic's abliterate() loop reaches into module.weight, and the fused block has no .weight → AttributeError: 'Qwen3_5MoeSparseMoeBlock' object has no attribute 'weight'.

The method

The patch (heretic-fused-experts.patch, included here) does three things:

  1. Surfaces the fused MoE block as one ablitable component (mlp.down_proj(fused)), excludes it from the LoRA/PEFT target set (it can't wrap a 3D Parameter), and abliterates both the routed experts and the dense shared_expert.

  2. Memory-safe reset via forward hooks. Abliterating down_proj by W -= λ·v(vᵀW) is mathematically identical to a rank-1 projection of the MoE block's output: y -= λ·v(vᵀy). So instead of editing (and backing up) the 32GB of 3D expert weights for every trial, a single forward hook per layer reproduces routed + shared expert ablation exactly, for any strength λ — at ~0.7 MB of state. Reset = remove the hook. This is what makes Heretic's per-trial Optuna search over a 256-expert model tractable without OOM. The chosen direction/strength is baked into the weights once, at save time.

  3. Respects the hybrid layers. The 30 linear-attention (Mamba/GDN) layers are left untouched; only attention o_proj + the MoE block are abliterated.

It also adds the if component == "mlp.down_proj(fused)": continue guard in abliterate() so the fused block is handled solely by the hooks (fixing the AttributeError).

Apply

The patch targets Heretic v1.3.0 (heretic/{model.py,main.py}):

# inside your Heretic install/container, from the package root:
patch -p1 < heretic-fused-experts.patch
# then run Heretic normally on a qwen3_5_moe model

Then run the standard Heretic flow. Watch the log for Abliterable components: ... mlp.down_proj(fused): N modules — if you only see attn.o_proj, the patch didn't take.

Results

On Qwen3.6-35B-A3B: refusals 86 → 5 / 100 at KL 0.0076 to the base; a 30-probe retention suite matched the base on every dimension across N=10 runs (extraction, multi-hop, reasoning, arithmetic, factual, code, language, instruction, format). Weights: Ember (BF16), Cinder (NVFP4).

Attribution & license

  • Method + patch: SparkyForge.
  • Built on: Heretic by Philipp Emanuel Weidmann — apply on top of Heretic v1.3.0; the patched portions are subject to Heretic's upstream license (see its repo).
  • Model: Qwen/Qwen3.6-35B-A3B (Apache 2.0), © the Qwen team.

Independent community work; not affiliated with NVIDIA, the Apache Software Foundation, the Qwen team, or the Heretic project.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-23method writeup394f3dd4.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration