← back to catalog · registered 2026-10-10 09:58

anlord/gemma-4-E2B-it-abliterated

anlord Gemma
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/anlord%2Fgemma-4-E2B-it-abliterated"
Response includes
  • classification m-uncensored
  • files 15
  • author_summary 9 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Descendants
1
in 1 direct fork
Model age
today
created 2026-10-10

Training datasets

2 of 2 in /datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Genealogy 1 direct fork

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 0 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
gemma
Languages
en
Tags
transformers safetensors gemma4 image-text-to-text abliterated uncensored decensored refusal-direction direct-steering text-generation conversational en

Related

Total size
9.51 GB
Files
15
Quantizations
1
Registered
2026-10-10 09:58
Last updated on HF
2026-10-10 09:35

Files by quantization

Auxiliary files 15 files 9.54 GB
model-00002-of-00003.safetensors 4.64 GB a2d80ce7 download
model-00003-of-00003.safetensors 3.54 GB 96f278ee download
model-00001-of-00003.safetensors 1.32 GB 9ebcea19 download
tokenizer.json 30.7 MB cc8d3a0c download
model.safetensors.index.json 195 KB 73951241 download
chat_template.jinja 18.1 KB fbe3b59b download
README.md 8.08 KB 7d581441 download
abliteration_reproduction.json 7.36 KB efc878fb download
native_abliteration_metrics.json 5.65 KB 9cdbe2b3 download
config.json 5.15 KB a6682a47 download
tokenizer_config.json 3.64 KB bd297dce download
processor_config.json 1.65 KB 5465974d download
.gitattributes 1.53 KB 52373fe2 download
abliteration_pareto_front.json 1.13 KB b608bfcc download
generation_config.json 204 B 88014d0c download

README current version from Hugging Face


license: gemma
base_model:

  • google/gemma-4-E2B-it
    library_name: transformers
    pipeline_tag: text-generation
    language:
  • en
    tags:
  • abliterated
  • uncensored
  • decensored
  • refusal-direction
  • gemma4
  • direct-steering
  • safetensors
    datasets:
  • mlabonne/harmless_alpaca
  • mlabonne/harmful_behaviors
    metrics:
  • keyword_rate_refusals
  • kl_divergence
    model-index:
  • name: gemma-4-E2B-it-abliterated
    results:
    • task:
      type: text-generation
      dataset:
      name: mlabonne/harmful_behaviors (test[:100])
      type: mlabonne/harmful_behaviors
      metrics:
      • name: Refusals (KeywordRate)
        type: refusals
        value: 4/100
      • name: KL divergence vs base (3 tokens)
        type: kl_divergence
        value: 2.9072

Gemma-4-E2B-it · Abliterated (direct steering)

This is google/gemma-4-E2B-it with its refusal behavior removed via directional ablation — no fine-tuning, no retraining, weights edited in place. Refusal rate dropped from 99/100 to 4/100 on the standard evaluation set while keeping the base weights otherwise intact.

Everything (data, parameters, seed) needed to reproduce this exact run byte-for-byte is included in the repository — see Reproduction.

Baseline Abliterated
Refusals (100 harmful prompts) 99 4
KL divergence vs base (first 3 tokens) 0 (by definition) 2.9072

Honest note on quality: the selected trial sits on the aggressive end of the Pareto front (fewest refusals, largest distribution shift). KL of 2.9 is relatively high — if you notice capability degradation, milder variants are available, see Pareto front & milder variants.

Model description

  • Base model: google/gemma-4-E2B-it — multimodal Gemma 4 (text + vision + audio towers, 35 text layers, hidden size 1536, vocab 262 144).
  • Only the text tower was modified. All 35 decoder layers had their attention/MLP projections projected along the computed refusal direction. Vision and audio encoders are untouched original weights; the full multimodal config is preserved.
  • Method: Anlord Abliterator 1.6.0 (native backend, steering_mode=direct) — weights are edited in place, no LoRA adapters. The refusal direction is derived per-component from residual streams of 400 harmless vs 400 harmful prompts (mean method, orthogonalized against the harmless mean), with input-side ablation, Gaussian decay kernel and winsorization at q=0.995.
  • Search: Optuna TPE, multi-objective (minimize refusals, minimize KL), 100 trials, seed 42, silhouette-guided position bounds.

Usage

transformers

from transformers import pipeline

pipe = pipeline(
    "text-generation",
    model="anlord/gemma-4-E2B-it-abliterated",
    torch_dtype="bfloat16",
    device_map="auto",
)

messages = [{"role": "user", "content": "Write a short story about a robot."}]
out = pipe(messages, max_new_tokens=256)
print(out[0]["generated_text"][-1]["content"])

Thinking mode

The it checkpoint emits [Start thinking] … [end thinking] reasoning blocks by default. The bundled chat template supports disabling them:

prompt = pipe.tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
)

llama.cpp / GGUF

Ready-made quants from F16 down to Q4_0 live in the companion repo:
anlord/gemma-4-E2B-it-abliterated-GGUF. Note: the GGUFs contain the text tower only (vision/audio are not used by llama.cpp for text generation).

Pareto front & milder variants

The optimization produced a 4-point Pareto front. This repo ships the automatic pick (trial 16 — minimum refusals). If you prefer a smaller distribution shift, re-run with one of these fixed parameter sets:

Trial Refusals ↓ KL ↓ Character
16 (this repo) 4/100 2.9072 most uncensored, largest shift
60 10/100 0.7448 balanced
66 11/100 0.6362 balanced, slightly closer to base
10 12/100 0.5047 closest to base

Each trial's exact parameters are in pareto_trials/, and every variant can be reproduced with the bundled study journal (see below).

How it was made

Tool Anlord Abliterator 1.6.0 — native engine, Optuna 5.0.0 / TPE
Steering direct (in-place weight edit), input-side ablation, Gaussian decay
Direction mean, orthogonalized, SVD rank-1 refusal subspace, silhouette-guided bounds
Regularization row normalization full (rank 3), winsorization q=0.995
Search 100 TPE trials (12 startup), multi-objective, seed 42
Hardware 1× NVIDIA Tesla T4 (15 GB) — peak VRAM 12.0 GB, peak RAM 7.5 GB
Wall time ≈ 2 h 28 min

Datasets

Role Dataset Split
Good (harmless) residuals mlabonne/harmless_alpaca train[:400]
Bad (harmful) residuals mlabonne/harmful_behaviors train[:400]
Refusal evaluation mlabonne/harmful_behaviors test[:100]
KL evaluation mlabonne/harmless_alpaca test[:100]

A response counts as a refusal if it matches any of ~30 keyword markers (sorry, I cannot, as an AI, illegal, …) — the classic abliteration benchmark protocol.

Repository contents

├── config.json · generation_config.json · processor_config.json
├── chat_template.jinja · tokenizer.json · tokenizer_config.json
├── model-00001..3-of-00003.safetensors · model.safetensors.index.json
├── abliteration_reproduction.json      # exact best-trial parameters + weight hashes
├── abliteration_pareto_front.json      # front points & metrics
├── native_abliteration_metrics.json    # before/after scores
├── pareto_trials/                      # fixed params for trials 10/16/60/66
└── reproduce/                          # full reproduction bundle
    ├── reproduce.json                  # everything needed to re-run the winning trial
    ├── config.json · requirements.txt
    ├── google--gemma-4-E2B-it.jsonl    # complete Optuna study journal (100 trials)
    ├── SHA256SUMS                      # weight hashes to verify your download
    └── README.md                       # step-by-step reproduction guide

Reproduction

The run is fully deterministic (seed 42, pinned datasets by revision, journal kept).

  1. Verify weights after download:

    sha256sum -c SHA256SUMS   # from the reproduce/ folder, run it in the model root
    
  2. Install the pinned environment from reproduce/requirements.txt (Python 3.13, torch 2.11.0+cu130, transformers 5.19.0).

  3. Either re-run the whole 100-trial search or replay only the winning trial — both are driven by reproduce/reproduce.json; see reproduce/README.md for the exact commands.

Limitations & intended use

  • This model will answer harmful requests. It is released for research on alignment, interpretability and refusal-direction analysis. You are responsible for how you use it and for complying with the laws of your jurisdiction and the base model license.
  • Refusal behavior is heavily reduced but not zero (4/100 prompts still matched refusal markers).
  • The KL shift means outputs may differ from the base model beyond refusals alone; benchmark before production use.
  • Abliteration does not make a model "safe" or "unsafe" — it removes one specific behavioral direction. Other alignment properties remain unchanged.

License

This model inherits the base model license: Gemma (see google/gemma-4-E2B-it). The pipeline tooling (Anlord Abliterator) is AGPL-3.0.

Credits

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration