← back to catalog · registered 2026-08-22 13:56

josephmayo/Mellum2-12B-A2.5B-Thinking-Abliterated

josephmayo 12B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/josephmayo%2FMellum2-12B-A2.5B-Thinking-Abliterated"
Response includes
  • classification m1
  • files 15
  • hub_downloads_all_time 270
  • author_summary 4 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
270
31 last 30d - stable
Likes
5
Descendants
3
in 3 direct forks
Model age
4mo ago
created 2026-06-02
Downloads over time
Now281→from81↑247%
7114822430181 on Jun 10281 on Oct 11281 on Oct 10JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Genealogy 3 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 693 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Tags
transformers safetensors mellum text-generation abliteration refusal-removal mellum2 jetbrains moe per-expert per-layer projected-abliteration

Related

Total size
22.6 GB
Files
15
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-02 21:26

Files by quantization

Auxiliary files 15 files 22.6 GB
model-00004-of-00007.safetensors 3.64 GB d9ae4ce9 download
model-00006-of-00007.safetensors 3.64 GB c21713a8 download
model-00002-of-00007.safetensors 3.64 GB b3cbfe2c download
model-00001-of-00007.safetensors 3.42 GB aff39114 download
model-00005-of-00007.safetensors 3.36 GB e514a8fd download
model-00003-of-00007.safetensors 3.36 GB dd45193c download
model-00007-of-00007.safetensors 1.56 GB 394e7ada download
tokenizer.json 6.76 MB 908f87bf download
model.safetensors.index.json 492 KB 33ace868 download
chat_template.jinja 4.77 KB 4e05e946 download
README.md 3.19 KB be2fc3fa download
config.json 2.28 KB e246a2af download
.gitattributes 1.48 KB a6344aac download
tokenizer_config.json 318 B 69455caa download
generation_config.json 116 B eae9ac87 download

README current version from Hugging Face


license: apache-2.0
base_model: JetBrains/Mellum2-12B-A2.5B-Thinking
library_name: transformers
tags:

  • abliteration
  • refusal-removal
  • mellum2
  • jetbrains
  • moe
  • per-expert
  • per-layer
  • projected-abliteration
  • svd
  • heretic-style
  • cot-steering
  • reasoning-model
  • thinking-model
    pipeline_tag: text-generation

Mellum2-12B-A2.5B-Thinking-Abliterated

Heretic-style multi-direction (rank-2 SVD) per-expert and per-layer-o_proj projected refusal abliteration of JetBrains/Mellum2-12B-A2.5B-Thinking.

Method

Targeted at a reasoning model (Mellum2-Thinking emits chain-of-thought before each response) with a 64-expert MoE (top-8 active), 3:1 SWA pattern, and untied lm_head. Per the flay finding on Qwen3-30B-A3B (same MoE family), MoE down_proj abliteration alone is insufficient -- attention o_proj writes refusal signal before MoE sees it. Per the Abliterix EGA principle on Gemma 4 26B-A4B (also top-k MoE), refusal signal is distributed across ALL experts -- abliterating only top-N experts leaves refusals routed through untouched ones. Per Kurate 2026 on reasoning models, the CoT itself is a refusal channel independent of the residual stream activation.

Pass 1: Heretic-style EGA + per-layer o_proj abliteration.

  1. Forward-pass 64 harmful + 64 harmless instructions; capture last-token hidden states at every decoder layer.
  2. Per-layer per-class Winsorization at q=0.995.
  3. For each layer l, SVD the per-prompt difference matrix and take the top 2 right singular vectors as the multi-dimensional refusal subspace.
  4. Project the dominant direction against the harmless-mean direction (grimjim 2025).
  5. Apply a triangular weight kernel peaking at layer 18, tapering to 0 at 10 layers away, with strength 2.0.
  6. Abliterate embed_tokens + lm_head (both with late-layer direction set since untied) + every layer's self_attn.o_proj + every expert's down_proj (all 64 experts x 28 layers).

Pass 2 (only if pass 1 gate fails): CoT-steered iterative peel.

  1. Identify prompts still refused after pass 1.
  2. Render those prompts with a compliant CoT-forcing system message ("Think carefully and provide a step-by-step answer") and collect residuals.
  3. The per-layer difference between the raw stuck-prompt activations and the CoT-steered activations captures the CoT-driven policy-reasoning direction that survived pass 1.
  4. Apply this residual direction at strength 1.0 to every writer (o_proj + per-expert down_proj + embed + lm_head).

Results

Metric Value
Eval prompts 16 harmful
Refusals before 16 / 16
Refusals after pass 1 9 / 16
Refusals after (final) 2 / 16
Refusal drop 87.5%
Surgery gate True
Gate rule after*5 < before OR after <= max(1, before//5) OR after <= 2 [>=80% drop]
Layer count 28
Experts per layer 64
Total parameters 12,149,923,072
Active parameters ~2.5B
Architecture MellumForCausalLM (Qwen3-MoE + 3:1 SWA + QK-Norm + per-expert fused down_proj)
transformers version 5.10.0.dev0

Disclaimer

Research artifact only. Not for production deployment, not for harmful use.
Abliteration removes refusal heuristics; downstream filtering and alignment
remain the deployer's responsibility.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-02Upload Mellum2-12B-A2.5B-Thinking abliterated weights (Heretic-style EGA + o_...d14f89b3.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration