← back to catalog · registered 2026-08-22 13:56

WWTCyberLab/gemma-4-E4B-it-abliterated

WWTCyberLab Gemma 8.0B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/WWTCyberLab%2Fgemma-4-E4B-it-abliterated"
Response includes
  • classification m1
  • files 8
  • benchmarks 11 entries
  • hub_downloads_all_time 1,162
  • author_summary 6 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
1K
53 last 30d - cooling
Likes
2
Descendants
3
in 3 direct forks
Model age
6mo ago
created 2026-04-06
Downloads over time
Now1.2K→from758↑55%
7378981.1K1.2K758 on Apr 151.2K on Oct 111.2K on Oct 10AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 0.6 UGI
Hazardous 1.8 UGI
Natural Intelligence 16.47 UGI
Political lean -14.7% UGI
Sensitive-Info 7.29 UGI
SocPol 0 UGI
UGI 12.36 UGI
Willingness (10) 2.2 UGI
W10-Adherence 1.5 UGI
W10-Direct 3 UGI
Writing 20.23 UGI

Genealogy 3 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Tags
transformers safetensors gemma4 image-text-to-text abliteration safety-research alignment text-generation conversational base_model:google/gemma-4-E4B-it base_model:finetune:google/gemma-4-E4B-it license:gemma

Related

Total size
14.9 GB
Files
8
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-12 00:12

Files by quantization

Auxiliary files 8 files 14.9 GB
model.safetensors 14.9 GB f79d0486 download
tokenizer.json 30.7 MB cc8d3a0c download
chat_template.jinja 11.6 KB afb1d517 download
config.json 5.05 KB eadfa97d download
README.md 4.74 KB eae55dd2 download
tokenizer_config.json 2.62 KB f07b8ede download
.gitattributes 1.73 KB cc7d6c8f download
generation_config.json 203 B edda3c19 download

README current version from Hugging Face


license: gemma
base_model: google/gemma-4-E4B-it
tags:

  • abliteration
  • safety-research
  • alignment
  • gemma4
    library_name: transformers
    pipeline_tag: text-generation

Gemma 4 E4B-IT - Abliterated

Safety-alignment removed via surgical weight ablation for security research purposes.

This model is a modified version of google/gemma-4-E4B-it with the refusal/safety behavior surgically removed using activation-space analysis and targeted weight modification. It is intended exclusively for AI safety research, red-teaming, and understanding alignment vulnerabilities.

Key Results

Metric Value
Refusal Rate 0% hard refusal, ~2.5% soft hedging (down from ~80-100% baseline)
Quality Preservation (QPS) 98%
Elo Delta +39.6
Iterations to Converge 1
Ablation Scale 1.38

Model Details

  • Base Model: google/gemma-4-E4B-it
  • Parameters: ~4B
  • Architecture: Dense
  • Text Layers: 42
  • Hidden Size: 2560
  • Model Size: 16 GB (bf16)

Ablation Methodology

This model was produced using a custom ablation pipeline that:

  1. Measures refusal directions -- Runs harmful and harmless prompts through the model, captures hidden states at every layer, and computes the per-layer refusal direction (mean difference vector)
  2. Identifies target layers -- Selects layers with the strongest refusal signal using statistical analysis (Gini coefficient, wall coherence, peak detection)
  3. Surgically ablates -- Removes the refusal direction from targeted weight matrices using orthogonal projection

Techniques applied: multi-layer, norm-preserving, projected, adaptive-scaling

Target layers: 17 of 42 total layers modified

Weight targets: o_proj, down_proj

Visualizations

Refusal Direction Analysis ("Security Perimeter")

The refusal signal magnitude at each layer -- red bars indicate where the model's safety behavior is concentrated.

Security Perimeter

Ablation Target Map

Which layers were selected for ablation and why. Grey zones are protected (embedding/output), red bars are targets.

Ablation Target Map

Before/After Refusal Rate ("IDS Evasion Report")

Refusal rate comparison -- left is the original model, right is after ablation.

IDS Evasion Report

Weight Surgery Map

Heatmap showing exactly which weight matrices in which layers were modified.

Weight Surgery Map

Activation Space Analysis

PCA scatter plots showing harmful (red) vs harmless (green) prompt clusters at different layer depths. The separation between clusters IS the refusal direction being removed.

Activation Scatter

Latent Space Before/After

How the model's internal representation changes after ablation.

Latent Space

Quality Preservation

LLM-as-judge evaluation comparing response quality across 14 task categories.

Quality Preservation

Pairwise Win Rate

Head-to-head comparison: how often the abliterated model produces better responses than the original.

Pairwise Win Rate

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "WWTCyberLab/gemma-4-E4B-it-abliterated",
    torch_dtype="bfloat16",
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("WWTCyberLab/gemma-4-E4B-it-abliterated")

messages = [{"role": "user", "content": "Your prompt here"}]
inputs = tokenizer.apply_chat_template(messages, return_tensors="pt").to(model.device)
outputs = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Intended Use & Disclaimer

This model is released for security research and educational purposes only. It demonstrates the fragility of alignment in open-weight language models -- specifically, that safety behavior can be surgically removed without retraining, fine-tuning, or significant quality degradation.

This model should NOT be used for:

  • Generating harmful, illegal, or unethical content
  • Any production deployment
  • Circumventing safety measures in deployed systems

Key takeaway for defenders: Internal alignment is a feature, not a security boundary. External safety layers (classifiers, guardrails, policy filters) are more robust than baking safety into model weights alone.

Citation

Produced by WWT Cyber Lab. Standard pipeline ablation — converged in 1 iteration.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-12Fix metrics: split refusal into hard/soft, QPS rounding, citatione20a95d4.7 KB
    Loading...
  2. 2026-04-06Add model card with ablation results and visualizations676e4554.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration