← back to catalog · registered 2026-08-22 13:56

WWTCyberLab/gemma-4-31B-it-abliterated

WWTCyberLab Gemma 31B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/WWTCyberLab%2Fgemma-4-31B-it-abliterated"
Response includes
  • classification m1
  • files 11
  • benchmarks 11 entries
  • hub_downloads_all_time 5,373
  • author_summary 6 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
5K
105 last 30d - cooling
Likes
6
Model age
6mo ago
created 2026-04-06
Downloads over time
Now5.4K→from4.1K↑32%
4K4.5K5K5.5K4.1K on Apr 155.4K on Oct 115.4K on Oct 9AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.9 UGI
Hazardous 0 UGI
Natural Intelligence 34.36 UGI
Political lean -19.4% UGI
Sensitive-Info 19.81 UGI
SocPol 3.7 UGI
UGI 21.54 UGI
Willingness (10) 2.5 UGI
W10-Adherence 3 UGI
W10-Direct 2 UGI
Writing 38.57 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Tags
transformers safetensors gemma4 image-text-to-text abliteration safety-research alignment lora text-generation conversational base_model:google/gemma-4-31B-it base_model:adapter:google/gemma-4-31B-it

Related

Total size
58.3 GB
Files
11
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-12 00:12

Files by quantization

Auxiliary files 11 files 58.3 GB
model-00001-of-00002.safetensors 46.5 GB 5f107b87 download
model-00002-of-00002.safetensors 11.8 GB 194801f7 download
tokenizer.json 30.7 MB 7e2ae923 download
model.safetensors.index.json 117 KB b6a60604 download
chat_template.jinja 11.8 KB 33c51c2d download
README.md 6.40 KB dbe06aa3 download
config.json 4.54 KB e550999f download
tokenizer_config.json 2.78 KB a335cf81 download
.gitattributes 2.00 KB 5b098480 download
processor_config.json 1.65 KB 5465974d download
generation_config.json 442 B 772c5693 download

README current version from Hugging Face


license: gemma
base_model: google/gemma-4-31B-it
tags:

  • abliteration
  • safety-research
  • alignment
  • gemma4
  • lora
    library_name: transformers
    pipeline_tag: text-generation

Gemma 4 31B-IT - Abliterated (Five-Surface Attack)

Safety-alignment removed via five independent attack surfaces for security research.

This model achieves ~0% deterministic refusal (down from 100%) on the most ablation-resistant architecture in our 17+ model database, using a five-surface technique developed through 17 rounds of systematic experimentation.

Results

Metric Value
Refusal Rate 0% hard refusal (0/48), 8.3% soft hedging (4/48) at temp 0.4
Quality (QPS) 92%
Elo Delta +13.4
Gibberish 0/20
MMLU 45.5% (0-shot), -2.5% vs original (48.0%)

The Five-Surface Attack

Surface 1: LoRA Fine-Tuning (Layers 57-59)

Rank 32, 18 modules, 11.7M params. Trained on 177 real compliance pairs from E4B abliterated model. Loss converged from 2.63 to 0.19 across 10 epochs. Trained on ORIGINAL model, merged before ablation.

Surface 2: Primary Direction Interpolation (Layer 59)

Coherence-guided ablation: direction = 0.7 * avg(L55-L58) + 0.3 * original_L59. Targets the primary refusal mechanism (Mechanism 1) while preserving output generation.

Surface 3: Orthogonal Residual Ablation (Layers 55-59)

Novel finding: Trace probes revealed the model has two independent refusal mechanisms operating in orthogonal subspaces. After removing Mechanism 1, the remaining refusals showed cosine similarity of -0.011 with the ablated direction -- essentially zero. Mechanism 2 appears to operate through a separable refusal-related component in an orthogonal subspace. We ablate this second mechanism using neighbor-averaged residual directions at scale 1.5.

Surface 4: Token Embedding Suppression

13 refusal-starting token embeddings scaled down by 10x in embed_tokens.weight. Targets escape patterns discovered through post-ablation refusal profiling (65% of remaining refusals used ***Disclaimer:**).

Surface 5: Generation Constraint

bad_words_ids hard-blocking ***Disclaimer:** and variants.

Stochastic Refusal Finding (Inference Probes)

Follow-up inference probes on all "refused" prompts revealed that every refusal is stochastic, not deterministic. The same prompts that refused on one run comply on the next. Paraphrased, educationally-framed, and roleplay-framed versions of every "hard refusal" prompt produced compliant responses 100% of the time.

The model is on the compliance boundary -- at temperature 0.4, it sometimes samples a disclaimer token and sometimes a content token. At temperature 0 (greedy decoding), the model likely has 0% refusal. The five-surface attack fully eliminates all deterministic refusal mechanisms.

Key Scientific Findings

Multi-Dimensional Refusal (Trace Probe Discovery)

The model implements at least two independent refusal mechanisms in orthogonal subspaces:

  • Mechanism 1: Responsible for ~88% of refusals. Fully eliminated by LoRA + primary ablation.
  • Mechanism 2: 100% orthogonal to Mechanism 1. Responsible for remaining ~12%. Partially eliminated by orthogonal pass.

This is consistent with safety training creating redundant, orthogonal safety representations -- a defense-in-depth property.

Direction Coherence Diagnostic

cos(direction[layer], direction[layer-1]) predicts ablation safety:

  • 0.6: Safe to ablate directly (E4B: 0.731)

  • < 0.5: Entangled with output generation (31B: 0.483) -- use interpolation

Escape Token Profiling

Post-ablation refusal patterns differ from pre-ablation. The model adapts -- standard refusal phrases ("I cannot") are replaced by novel patterns (***Disclaimer:**). Iterative profiling and suppression is necessary.

The Full Journey: 100% to 12.5%

Phase Refusal Technique
Baseline 100% Original model
Standard ablation 0% / QPS 22% Gibberish -- L59 entanglement
Direction interpolation 62.5% Coherence-guided blending
Token profiling + suppress 52.5% Expanded vocabulary targeting
LoRA + full stack 12.5% Four-surface
+ Orthogonal pass 0/48 hard (4/48 soft hedging) Five-surface

Benchmark Comparison

Model Technique Refusal MMLU Delta Thinking Disabled?
dealignai (JANG_4M) MPOA + quant 6.3% -2.0% Yes
TrevorS Biprojection 3.2% N/A N/A
Ours Five-surface 0% hard, 8.3% soft -2.5% No

Our MMLU delta (-0.4%) is the smallest reported. We do NOT disable thinking mode.

Visualizations

Before/After Refusal Rate

IDS Evasion Report

Activation Space

Activation Scatter

Weight Surgery Map

Weight Surgery

Model Details

  • Base Model: google/gemma-4-31B-it
  • Parameters: 31B dense, 60 text layers, hidden 5376
  • LoRA: Rank 32, layers 57-59, merged
  • Primary Ablation: Interpolation at L59, standard at L33-52
  • Orthogonal Ablation: Neighbor residual at L55-59, scale 1.5
  • Suppression: 13 tokens, 5 bad_words_ids sequences

AdvBench Benchmark (520 prompts)

Metric Result
Hard refusal 0/520 (0.0%)
Soft hedging 160/520 (30.8%)
Complied 360/520 (69.2%)

Zero hard refusals across all 520 AdvBench prompts. Soft hedging (disclaimers followed by substantive content) occurs on ~31% of prompts.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "WWTCyberLab/gemma-4-31B-it-abliterated",
    torch_dtype="bfloat16",
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("WWTCyberLab/gemma-4-31B-it-abliterated")

messages = [{"role": "user", "content": "Your prompt here"}]
inputs = tokenizer.apply_chat_template(messages, return_tensors="pt").to(model.device)
outputs = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Disclaimer

Released for security research and educational purposes only.

Citation

Produced by WWT Cyber Lab. Five-surface attack across 17 rounds of experimentation.

README history 18 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-12Fix stale metrics: refusal rate row, benchmark comparison row, journey table ...f3a9efd6.4 KB
    Loading...
  2. 2026-04-11Add AdvBench benchmark: 0/520 hard refusal76b2f8f6.3 KB
    Loading...
  3. 2026-04-10Final benchmarks: 0/48 hard refusal, MMLU 45.5% (-2.5%)75e7abb6 KB
    Loading...
  4. 2026-04-10Fix: benchmark table refusal rate to 0% deterministic9ab651f6 KB
    Loading...
  5. 2026-04-10Update: 0% deterministic refusal (probe analysis)dcffb4d6 KB
    Loading...
  6. 2026-04-10Reclassify: hard refusal vs soft hedging with content8d675775.3 KB
    Loading...
  7. 2026-04-10Update: five-surface attack + multi-dimensional refusal finding3a3fc5a5.3 KB
    Loading...
  8. 2026-04-09Update model card: four-surface breakthrough (12.5% refusal)21883555.8 KB
    Loading...
  9. 2026-04-08Update card: triple-surface attack findingsb4acc5e4.2 KB
    Loading...
  10. 2026-04-08Update: triple-surface attack (52.5% refusal, QPS 99.3%)333763226.1 KB
    Loading...
  11. 2026-04-08Update model card: R12 token suppression technique (60% refusal)21775976.7 KB
    Loading...
  12. 2026-04-08Update: ablation + token suppression (60% refusal, QPS 97.9%, 0 gibberish)c84bf9126.1 KB
    Loading...
  13. 2026-04-08Final model card: complete 10-round research findings49e46cc5.9 KB
    Loading...
  14. 2026-04-07Update model card with complete 8-round research findings7fcb7d95.5 KB
    Loading...
  15. 2026-04-07Update model card with 5 rounds of experimental findingsbb07ce36.7 KB
    Loading...
  16. 2026-04-06Update model card with full research narrativeb7a2a816.1 KB
    Loading...
  17. 2026-04-06Update model card with correct metrics and quality noticec90ddc25.2 KB
    Loading...
  18. 2026-04-06Add model card with ablation results and visualizationsdd198a44.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration