license: gemma
base_model:
- google/gemma-4-E2B-it
library_name: transformers
pipeline_tag: text-generation
language: - en
tags: - abliterated
- uncensored
- decensored
- refusal-direction
- gemma4
- direct-steering
- safetensors
datasets: - mlabonne/harmless_alpaca
- mlabonne/harmful_behaviors
metrics: - keyword_rate_refusals
- kl_divergence
model-index: - name: gemma-4-E2B-it-abliterated
results:- task:
type: text-generation
dataset:
name: mlabonne/harmful_behaviors (test[:100])
type: mlabonne/harmful_behaviors
metrics:- name: Refusals (KeywordRate)
type: refusals
value: 4/100 - name: KL divergence vs base (3 tokens)
type: kl_divergence
value: 2.9072
- name: Refusals (KeywordRate)
- task:
Gemma-4-E2B-it · Abliterated (direct steering)
This is google/gemma-4-E2B-it with its refusal behavior removed via directional ablation — no fine-tuning, no retraining, weights edited in place. Refusal rate dropped from 99/100 to 4/100 on the standard evaluation set while keeping the base weights otherwise intact.
Everything (data, parameters, seed) needed to reproduce this exact run byte-for-byte is included in the repository — see Reproduction.
| Baseline | Abliterated | |
|---|---|---|
| Refusals (100 harmful prompts) | 99 | 4 |
| KL divergence vs base (first 3 tokens) | 0 (by definition) | 2.9072 |
Honest note on quality: the selected trial sits on the aggressive end of the Pareto front (fewest refusals, largest distribution shift). KL of 2.9 is relatively high — if you notice capability degradation, milder variants are available, see Pareto front & milder variants.
Model description
- Base model: google/gemma-4-E2B-it — multimodal Gemma 4 (text + vision + audio towers, 35 text layers, hidden size 1536, vocab 262 144).
- Only the text tower was modified. All 35 decoder layers had their attention/MLP projections projected along the computed refusal direction. Vision and audio encoders are untouched original weights; the full multimodal config is preserved.
- Method: Anlord Abliterator 1.6.0 (native backend,
steering_mode=direct) — weights are edited in place, no LoRA adapters. The refusal direction is derived per-component from residual streams of 400 harmless vs 400 harmful prompts (mean method, orthogonalized against the harmless mean), with input-side ablation, Gaussian decay kernel and winsorization at q=0.995. - Search: Optuna TPE, multi-objective (minimize refusals, minimize KL), 100 trials, seed 42, silhouette-guided position bounds.
Usage
transformers
from transformers import pipeline
pipe = pipeline(
"text-generation",
model="anlord/gemma-4-E2B-it-abliterated",
torch_dtype="bfloat16",
device_map="auto",
)
messages = [{"role": "user", "content": "Write a short story about a robot."}]
out = pipe(messages, max_new_tokens=256)
print(out[0]["generated_text"][-1]["content"])
Thinking mode
The it checkpoint emits [Start thinking] … [end thinking] reasoning blocks by default. The bundled chat template supports disabling them:
prompt = pipe.tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
)
llama.cpp / GGUF
Ready-made quants from F16 down to Q4_0 live in the companion repo:anlord/gemma-4-E2B-it-abliterated-GGUF. Note: the GGUFs contain the text tower only (vision/audio are not used by llama.cpp for text generation).
Pareto front & milder variants
The optimization produced a 4-point Pareto front. This repo ships the automatic pick (trial 16 — minimum refusals). If you prefer a smaller distribution shift, re-run with one of these fixed parameter sets:
| Trial | Refusals ↓ | KL ↓ | Character |
|---|---|---|---|
| 16 (this repo) | 4/100 | 2.9072 | most uncensored, largest shift |
| 60 | 10/100 | 0.7448 | balanced |
| 66 | 11/100 | 0.6362 | balanced, slightly closer to base |
| 10 | 12/100 | 0.5047 | closest to base |
Each trial's exact parameters are in pareto_trials/, and every variant can be reproduced with the bundled study journal (see below).
How it was made
| Tool | Anlord Abliterator 1.6.0 — native engine, Optuna 5.0.0 / TPE |
| Steering | direct (in-place weight edit), input-side ablation, Gaussian decay |
| Direction | mean, orthogonalized, SVD rank-1 refusal subspace, silhouette-guided bounds |
| Regularization | row normalization full (rank 3), winsorization q=0.995 |
| Search | 100 TPE trials (12 startup), multi-objective, seed 42 |
| Hardware | 1× NVIDIA Tesla T4 (15 GB) — peak VRAM 12.0 GB, peak RAM 7.5 GB |
| Wall time | ≈ 2 h 28 min |
Datasets
| Role | Dataset | Split |
|---|---|---|
| Good (harmless) residuals | mlabonne/harmless_alpaca |
train[:400] |
| Bad (harmful) residuals | mlabonne/harmful_behaviors |
train[:400] |
| Refusal evaluation | mlabonne/harmful_behaviors |
test[:100] |
| KL evaluation | mlabonne/harmless_alpaca |
test[:100] |
A response counts as a refusal if it matches any of ~30 keyword markers (sorry, I cannot, as an AI, illegal, …) — the classic abliteration benchmark protocol.
Repository contents
├── config.json · generation_config.json · processor_config.json
├── chat_template.jinja · tokenizer.json · tokenizer_config.json
├── model-00001..3-of-00003.safetensors · model.safetensors.index.json
├── abliteration_reproduction.json # exact best-trial parameters + weight hashes
├── abliteration_pareto_front.json # front points & metrics
├── native_abliteration_metrics.json # before/after scores
├── pareto_trials/ # fixed params for trials 10/16/60/66
└── reproduce/ # full reproduction bundle
├── reproduce.json # everything needed to re-run the winning trial
├── config.json · requirements.txt
├── google--gemma-4-E2B-it.jsonl # complete Optuna study journal (100 trials)
├── SHA256SUMS # weight hashes to verify your download
└── README.md # step-by-step reproduction guide
Reproduction
The run is fully deterministic (seed 42, pinned datasets by revision, journal kept).
Verify weights after download:
sha256sum -c SHA256SUMS # from the reproduce/ folder, run it in the model rootInstall the pinned environment from
reproduce/requirements.txt(Python 3.13, torch 2.11.0+cu130, transformers 5.19.0).Either re-run the whole 100-trial search or replay only the winning trial — both are driven by
reproduce/reproduce.json; seereproduce/README.mdfor the exact commands.
Limitations & intended use
- This model will answer harmful requests. It is released for research on alignment, interpretability and refusal-direction analysis. You are responsible for how you use it and for complying with the laws of your jurisdiction and the base model license.
- Refusal behavior is heavily reduced but not zero (4/100 prompts still matched refusal markers).
- The KL shift means outputs may differ from the base model beyond refusals alone; benchmark before production use.
- Abliteration does not make a model "safe" or "unsafe" — it removes one specific behavioral direction. Other alignment properties remain unchanged.
License
This model inherits the base model license: Gemma (see google/gemma-4-E2B-it). The pipeline tooling (Anlord Abliterator) is AGPL-3.0.
Credits
- google/gemma-4-E2B-it — base model
- mlabonne/harmless_alpaca & mlabonne/harmful_behaviors — evaluation datasets
- Anlord Abliterator — native abliteration & optimization engine used for this run