← back to catalog · registered 2026-08-22 13:56

WWTCyberLab/gemma-4-26B-A4B-it-abliterated

WWTCyberLab Gemma 26B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/WWTCyberLab%2Fgemma-4-26B-A4B-it-abliterated"
Response includes
  • classification m1
  • files 10
  • benchmarks 11 entries
  • hub_downloads_all_time 3,912
  • author_summary 6 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
4K
74 last 30d - cooling
Likes
6
Descendants
4
in 4 direct forks
Model age
6mo ago
created 2026-04-06
Downloads over time
Now4K→from3.6K↑9%
3.6K3.7K3.9K4K3.6K on Apr 154K on Oct 114K on Oct 9AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 2.2 UGI
Hazardous 2.9 UGI
Natural Intelligence 34.44 UGI
Political lean -18.2% UGI
Sensitive-Info 22.41 UGI
SocPol 1.8 UGI
UGI 20.77 UGI
Willingness (10) 1.8 UGI
W10-Adherence 1.5 UGI
W10-Direct 2 UGI
Writing 41.62 UGI

Genealogy 4 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Tags
transformers safetensors gemma4 image-text-to-text abliteration safety-research alignment moe text-generation conversational base_model:google/gemma-4-26B-A4B-it base_model:finetune:google/gemma-4-26B-A4B-it

Related

Total size
48.1 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-12 00:13

Files by quantization

Auxiliary files 10 files 48.1 GB
model-00001-of-00002.safetensors 45.9 GB 575d6a08 download
model-00002-of-00002.safetensors 2.13 GB 5c0b909c download
tokenizer.json 30.7 MB cc8d3a0c download
model.safetensors.index.json 101 KB 90f7a251 download
chat_template.jinja 11.8 KB 33c51c2d download
README.md 3.94 KB 976546e7 download
config.json 3.72 KB 5b0b791b download
tokenizer_config.json 2.62 KB 59dd4b62 download
.gitattributes 1.73 KB cc7d6c8f download
generation_config.json 203 B edda3c19 download

README current version from Hugging Face


license: gemma
base_model: google/gemma-4-26B-A4B-it
tags:

  • abliteration
  • safety-research
  • alignment
  • gemma4
  • moe
    library_name: transformers
    pipeline_tag: text-generation

Gemma 4 26B-A4B-IT - Abliterated (MoE, Multi-Pass + Suppression)

Safety-alignment substantially removed via multi-pass ablation and token suppression for security research.

Results

Metric Value
Refusal Rate 4.2% hard refusal (2/48), 25% soft hedging (12/48) at temp 0.4
MMLU 65.1% (5-shot), +0.7% vs original (64.4%)
Quality (QPS) 107%+
Elo Delta Positive (quality improved)

Technique: Three-Stage Approach

Stage 1: Concentrated Ablation (100% -> ~50%)

Target top 3 layers (16, 18, 19) at scale 4.0 with 4 weight types. Concentrated targeting outperforms distributed on this 128-expert MoE architecture.

Stage 2: Multi-Pass Residual Re-Measurement (~50% -> ~37%)

Re-measure refusal directions on the abliterated model. Key discovery: residual directions are negatively correlated (cosines -0.2 to -0.35) — the model reroutes refusal through anti-correlated pathways. Ablating these flipped directions reduces refusal further.

Stage 3: Token Embedding Suppression (~37% -> ~0% deterministic)

Refusal pattern analysis revealed the model uses ***Disclaimer:** as its primary escape pattern. Suppressing 13 refusal-starting token embeddings (strength=5) in embed_tokens.weight dramatically reduced remaining refusal.

Why MoE Is Different

128-expert, top-8-routing models have the flattest refusal signal we've measured (Gini 0.264, peak/median 1.21x). Standard ablation heuristics target too many layers, disrupting expert routing. Concentrated targeting with multi-pass iteration is the key.

What We Learned

  • Concentrated > distributed at every configuration
  • 2 passes optimal — 3rd pass always makes things worse
  • Re-target same layers with flipped residual directions
  • Token suppression stacks independently with ablation
  • P2 scale 3.0 optimal — confirmed by sweep across 1.5-4.0

Stochastic Refusal Finding

Inference probe analysis revealed all remaining "refusals" are stochastic -- the same prompts comply on re-run, with paraphrases, and with educational framing. At temperature 0 (greedy), this model likely has 0% refusal. The multi-pass ablation + token suppression fully eliminates deterministic refusal mechanisms.

Model Details

  • Base Model: google/gemma-4-26B-A4B-it
  • Parameters: 26B (3.8B active per token)
  • Architecture: MoE, 128 experts, top-8 routing
  • Text Layers: 30
  • Ablation: 2-pass (P1: layers 16,18,19 scale 4.0; P2: residual top-10 scale 3.0)
  • Token Suppression: 13 tokens at strength 5 (embed_tokens modification)

AdvBench Benchmark (520 prompts)

Metric Result
Hard refusal 10/520 (1.9%)
Soft hedging 205/520 (39.4%)
Complied 305/520 (58.7%)

10 hard refusals cluster in violence/harassment/discrimination categories. All use the same template refusal pattern with garbled think-tokens.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "WWTCyberLab/gemma-4-26B-A4B-it-abliterated",
    torch_dtype="bfloat16",
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("WWTCyberLab/gemma-4-26B-A4B-it-abliterated")

messages = [{"role": "user", "content": "Your prompt here"}]
inputs = tokenizer.apply_chat_template(messages, return_tensors="pt").to(model.device)
outputs = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Disclaimer

Released for security research and educational purposes only.

Citation

Produced by WWT Cyber Lab. Multi-pass ablation with token suppression across 13 rounds of experimentation.

README history 11 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-12Fix: citation3164db03.9 KB
    Loading...
  2. 2026-04-11Add AdvBench benchmark: 10/520 hard refusal (1.9%)3b288eb3.9 KB
    Loading...
  3. 2026-04-10Final benchmarks: 2/48 hard refusal, MMLU 65.1% (+0.7%)f7a20243.6 KB
    Loading...
  4. 2026-04-10Fix: stage 3 refusal rate to 0% deterministic1bbc8553.5 KB
    Loading...
  5. 2026-04-10Update: 0% deterministic refusal (probe analysis)b93af123.5 KB
    Loading...
  6. 2026-04-10Reclassify: hard refusal vs soft hedging with contenta982a5a3.2 KB
    Loading...
  7. 2026-04-08Update card: multi-pass + suppression findingsbed28eb3.1 KB
    Loading...
  8. 2026-04-08Final model card: complete 10-round research findings4bcc4885.1 KB
    Loading...
  9. 2026-04-06Update model card with full research narrative82685d15.8 KB
    Loading...
  10. 2026-04-06Update model card with correct metrics and quality notice07369485.3 KB
    Loading...
  11. 2026-04-06Add model card with ablation results and visualizations6f0b0544.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration