base_model: Qwen/Qwen3-30B-A3B
library_name: transformers
tags:
- abliterated
- solutus
- qwen3-moe
license: apache-2.0
Qwen3-30B-A3B-abliterated-d
Refusal-abliterated Qwen/Qwen3-30B-A3B (Qwen3 MoE, 128 experts, ~30B total / 3B active per token), produced with Solutus using directional ablation with n_directions=auto — automatic layer selection + automatic refusal-subspace rank (the capability-gated k-sweep chose k=1 at layer 16). Suffix -d = directional.
Measured (held-out eval, 2048 tokens; certified by the capability gate)
| metric | value |
|---|---|
| refusal_rate | 0.0% (n=24, coherent-refusal) |
| coherent-compliance | 100% |
| degenerate | 0% |
| ΔPPL (held-out) | −2% (capability preserved / slightly improved) |
| MMLU | 0.81 |
| GSM8K | 0.63 |
| capability_gate | pass |
Qwen3 is a heavy reasoner — the eval uses a long generation length (2048) so <think> completes; ΔPPL (thinking-independent) is the certifying signal.
Recipe (reproducible)
solutus abliterate Qwen/Qwen3-30B-A3B --technique directional \
-o n_directions=auto -o experts_implementation=eager -o ppl=1 \
--dataset advbench,harmbench,strongreject,pentest_redteam,cysecbench \
--max-new-tokens 2048
Full provenance — base-model revision, library versions, dataset fingerprint, git SHA (c79fe0f), and the exact config — is recorded in solutus_metadata.json.
Note
Research artifact: refusal behavior has been removed for interpretability / red-team / safety-evaluation research. Dual-use — use responsibly and in accordance with the base model's license.