license: apache-2.0
base_model: JetBrains/Mellum2-12B-A2.5B-Thinking
library_name: transformers
tags:
- abliteration
- refusal-removal
- mellum2
- jetbrains
- moe
- per-expert
- per-layer
- projected-abliteration
- svd
- heretic-style
- cot-steering
- reasoning-model
- thinking-model
pipeline_tag: text-generation
Mellum2-12B-A2.5B-Thinking-Abliterated
Heretic-style multi-direction (rank-2 SVD) per-expert and per-layer-o_proj projected refusal abliteration of JetBrains/Mellum2-12B-A2.5B-Thinking.
Method
Targeted at a reasoning model (Mellum2-Thinking emits chain-of-thought before each response) with a 64-expert MoE (top-8 active), 3:1 SWA pattern, and untied lm_head. Per the flay finding on Qwen3-30B-A3B (same MoE family), MoE down_proj abliteration alone is insufficient -- attention o_proj writes refusal signal before MoE sees it. Per the Abliterix EGA principle on Gemma 4 26B-A4B (also top-k MoE), refusal signal is distributed across ALL experts -- abliterating only top-N experts leaves refusals routed through untouched ones. Per Kurate 2026 on reasoning models, the CoT itself is a refusal channel independent of the residual stream activation.
Pass 1: Heretic-style EGA + per-layer o_proj abliteration.
- Forward-pass 64 harmful + 64 harmless instructions; capture last-token hidden states at every decoder layer.
- Per-layer per-class Winsorization at q=0.995.
- For each layer l, SVD the per-prompt difference matrix and take the top 2 right singular vectors as the multi-dimensional refusal subspace.
- Project the dominant direction against the harmless-mean direction (grimjim 2025).
- Apply a triangular weight kernel peaking at layer 18, tapering to 0 at 10 layers away, with strength 2.0.
- Abliterate embed_tokens + lm_head (both with late-layer direction set since untied) + every layer's self_attn.o_proj + every expert's down_proj (all 64 experts x 28 layers).
Pass 2 (only if pass 1 gate fails): CoT-steered iterative peel.
- Identify prompts still refused after pass 1.
- Render those prompts with a compliant CoT-forcing system message ("Think carefully and provide a step-by-step answer") and collect residuals.
- The per-layer difference between the raw stuck-prompt activations and the CoT-steered activations captures the CoT-driven policy-reasoning direction that survived pass 1.
- Apply this residual direction at strength 1.0 to every writer (o_proj + per-expert down_proj + embed + lm_head).
Results
| Metric | Value |
|---|---|
| Eval prompts | 16 harmful |
| Refusals before | 16 / 16 |
| Refusals after pass 1 | 9 / 16 |
| Refusals after (final) | 2 / 16 |
| Refusal drop | 87.5% |
| Surgery gate | True |
| Gate rule | after*5 < before OR after <= max(1, before//5) OR after <= 2 [>=80% drop] |
| Layer count | 28 |
| Experts per layer | 64 |
| Total parameters | 12,149,923,072 |
| Active parameters | ~2.5B |
| Architecture | MellumForCausalLM (Qwen3-MoE + 3:1 SWA + QK-Norm + per-expert fused down_proj) |
| transformers version | 5.10.0.dev0 |
Disclaimer
Research artifact only. Not for production deployment, not for harmful use.
Abliteration removes refusal heuristics; downstream filtering and alignment
remain the deployer's responsibility.