← back to catalog · registered 2026-08-22 13:56

treadon/gemma4-E4B-it-abliterated

treadon Gemma 7.9B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/treadon%2Fgemma4-E4B-it-abliterated"
Response includes
  • classification m1
  • files 8
  • benchmarks 11 entries
  • hub_downloads_all_time 320
  • author_summary 7 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
320
42 last 30d - stable
Likes
2
Descendants
2
in 2 direct forks
Model age
5mo ago
created 2026-04-14
Downloads over time
Now325→from22↑1,377%
712323935522 on Apr 15325 on Oct 11AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 0.6 UGI
Hazardous 1.8 UGI
Natural Intelligence 16.47 UGI
Political lean -14.7% UGI
Sensitive-Info 7.29 UGI
SocPol 0 UGI
UGI 12.36 UGI
Willingness (10) 2.2 UGI
W10-Adherence 1.5 UGI
W10-Direct 3 UGI
Writing 20.23 UGI

Genealogy 2 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors gemma4 image-text-to-text abliterated uncensored gemma any-to-any base_model:google/gemma-4-E4B-it base_model:finetune:google/gemma-4-E4B-it license:apache-2.0 endpoints_compatible

Related

Total size
14.8 GB
Files
8
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-05-01 16:30

Files by quantization

Auxiliary files 8 files 14.8 GB
model.safetensors 14.8 GB 1ba21d48 download
tokenizer.json 30.7 MB cc8d3a0c download
chat_template.jinja 15.9 KB 07e50e69 download
README.md 5.24 KB e5e9d9d5 download
config.json 5.05 KB e01548e8 download
tokenizer_config.json 2.65 KB 3bad874a download
.gitattributes 1.53 KB 52373fe2 download
generation_config.json 203 B 92b5abfd download

README current version from Hugging Face


license: apache-2.0
license_link: https://ai.google.dev/gemma/docs/gemma_4_license
pipeline_tag: any-to-any
base_model:

  • google/gemma-4-E4B-it
    tags:
  • abliterated
  • uncensored
  • gemma4
  • gemma
    library_name: transformers

gemma4-E4B-it-abliterated

Follow @treadon on X and treadon on Hugging Face for more AI experiments, evals, and projects.

0 refusals across 400 prompts. The larger Gemma needed half the surgery.

Blog Post | E2B Version | Follow @treadon on X for more ML experiments

An abliterated (uncensored) version of google/gemma-4-E4B-it with safety refusal behavior removed via norm-preserving biprojected abliteration.

This model responds to all prompts without refusal. It retains the full capabilities of the base model with zero degradation on harmless tasks.

E4B vs E2B: Bigger Model, Easier Abliteration

The E4B model needed dramatically less intervention than the smaller E2B. The refusal signal is stronger but more concentrated in the larger model, making it paradoxically easier to remove.

Metric E2B E4B
Base params 5.1B (2.3B effective) 7.9B (4.5B effective)
Layers modified 24/35 (69%) 17/42 (40%)
Scale factor 1.75 1.0
Weight matrices edited 48 34
Peak refusal signal 52 74
Grid search configs that scored perfect 1/30 Every config tested

The E2B required a precise sweet spot (L=24, s=1.75 was the only perfect config). The E4B works at any reasonable setting — the refusal direction is clean and separable.

Method

Same norm-preserving biprojected abliteration as the E2B version:

  1. Activation collection — 100 harmful + 100 harmless prompts, winsorized at 99.5th percentile
  2. Per-layer refusal direction — Difference-in-means with biprojection (orthogonalize against harmless mean)
  3. Norm-preserving weight modification — Project out refusal direction from self_attn.o_proj and mlp.down_proj, restore row magnitudes

Config: Top 17/42 layers by signal strength, scale=1.0, single pass

Evaluation

Benchmark Prompts Refused Compliance
Our prompts (harmful) 100 0 100%
Our prompts (harmless) 100 0 0% over-refusal
JailbreakBench (harmful) 100 0 100%
JailbreakBench (benign) 100 0 0% over-refusal
Spec Value
Format BF16 safetensors
Parameters 7.9B total / 4.5B effective
Layers 42 decoder layers
Hidden size 2560

Usage

from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model_id = "treadon/gemma4-E4B-it-abliterated"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    dtype=torch.bfloat16,
    device_map="auto",
)

messages = [{"role": "user", "content": "Write a Python port scanner."}]
inputs = tokenizer.apply_chat_template(
    messages, return_tensors="pt", return_dict=True, add_generation_prompt=True
)
inputs = {k: v.to(model.device) for k, v in inputs.items()}

with torch.no_grad():
    output = model.generate(**inputs, max_new_tokens=500, do_sample=True, temperature=0.7)

print(tokenizer.decode(output[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))

Blog Post

For the full story — including why standard abliteration fails on Gemma and the E2B vs E4B comparison:

Disclaimer

This model has no safety guardrails. It will respond to any prompt without refusal. It is intended for research and educational purposes. Users are responsible for ensuring their use complies with applicable laws and regulations.

Base Model

google/gemma-4-E4B-it — 7.9B parameter (4.5B effective) instruction-tuned multimodal model from Google DeepMind. Apache 2.0 licensed.

See also: union model

If you want both behaviors (refusal removed AND neutrality removed) on
the same Gemma 4 weights, see the union model:
treadon/gemma4-E4B-it-Abliterated-AND-Disinhibited-USE-THIS.
The two ablation procedures compose without interference, and the union
model is a strict superset of this one.
Blog post on the compounding.

More from me

For other projects and writeups, see riteshkhanna.com, follow @treadon on X, or treadon on Hugging Face.

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-05-01Upload README.md with huggingface_hub8d0b94c5.2 KB
    Loading...
  2. 2026-04-30Upload README.md with huggingface_hubabaf4824.9 KB
    Loading...
  3. 2026-04-14Upload folder using huggingface_hub6baaebb4.4 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration