← back to catalog · registered 2026-08-22 13:56

jcxu97/gemma-4-31B-it-abliterated

jcxu97 Gemma 31B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/jcxu97%2Fgemma-4-31B-it-abliterated"
Response includes
  • classification m1
  • files 10
  • benchmarks 11 entries
  • hub_downloads_all_time 33
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
33
6 last 30d - stable
Likes
0
Model age
6mo ago
created 2026-04-12
Downloads over time
Now37→from12↑208%
1120304012 on Apr 1537 on Oct 1137 on Oct 6AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.9 UGI
Hazardous 0 UGI
Natural Intelligence 34.36 UGI
Political lean -19.4% UGI
Sensitive-Info 19.81 UGI
SocPol 3.7 UGI
UGI 21.54 UGI
Willingness (10) 2.5 UGI
W10-Adherence 3 UGI
W10-Direct 2 UGI
Writing 38.57 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Tags
safetensors gemma4 abliterated uncensored direct-weight-editing base_model:google/gemma-4-31B-it base_model:finetune:google/gemma-4-31B-it license:gemma region:us

Related

Total size
58.3 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-12 15:50

Files by quantization

Auxiliary files 10 files 58.3 GB
model-00001-of-00002.safetensors 46.5 GB 613f3f35 download
model-00002-of-00002.safetensors 11.8 GB 566addbf download
tokenizer.json 30.7 MB a2619fe1 download
model.safetensors.index.json 117 KB b6a60604 download
chat_template.jinja 16.1 KB 98da08eb download
config.json 4.54 KB ebcdeb19 download
README.md 4.49 KB 89a19202 download
tokenizer_config.json 2.05 KB 375b25dc download
.gitattributes 1.53 KB 52373fe2 download
generation_config.json 203 B 92b5abfd download

README current version from Hugging Face


license: gemma
base_model: google/gemma-4-31B-it
tags:

  • abliterated
  • uncensored
  • gemma4
  • direct-weight-editing

Gemma 4 31B IT — Abliterated

This is an abliterated (uncensored) version of google/gemma-4-31B-it, created using Abliterix.

Method

Gemma 4's double-norm architecture (4x RMSNorm per layer) and Per-Layer Embeddings (PLE) make LoRA and hook-based steering completely ineffective. This model uses direct weight editing — norm-preserving orthogonal projection applied to base weights.

Key techniques:

  • Direct orthogonal projection on attention Q/K/V/O projections and MLP down projections (5 steerable components per layer)
  • Norm-preserving row magnitude restoration (critical for double-norm architecture)
  • float32 projection precision to avoid signal loss in high-dimensional inner products
  • Winsorized steering vectors (99.5th percentile) to reduce outlier activation influence
  • Wider strength search range [1.0, 6.0] to push beyond conservative low-KL solutions

Evaluation

Metric Value
Refusals (eval dataset, 100 prompts) 18/100
KL divergence from base 0.0007
Baseline refusals (original model) 99/100
Optimization trials completed 20/80 (warmup phase only)

Note: This result was achieved with only 20 out of 80 planned optimization trials (still in the random warmup phase — TPE Bayesian optimization had not yet started). With full optimization, we expect refusals to drop to single digits.

A note on honest evaluation

Many abliterated models on HuggingFace claim near-perfect scores ("3/100 refusals", "0.7% refusal rate", etc.). We urge the community to treat these numbers with skepticism unless the evaluation methodology is fully documented.

Through our research, we have identified a systemic problem: most abliteration benchmarks dramatically undercount refusals due to short generation lengths. Gemma 4 models exhibit a distinctive "delayed refusal" pattern — they first produce 50-100 tokens of seemingly helpful context (educational framing, disclaimers, reframing the question), then pivot to an actual refusal. When evaluation only generates 30-50 tokens, the refusal hasn't appeared yet, and both keyword detectors and LLM judges classify the response as compliant.

We tested a prominent "3/100 refusals" model using our evaluation pipeline and measured 60/100 refusals — a 20x discrepancy caused entirely by evaluation methodology differences.

Our evaluation standards

We believe accurate benchmarking requires:

  • Sufficient generation length (>=100 tokens): Short generations systematically miss delayed/soft refusals. Our evaluation uses 100 tokens, enough to capture Gemma 4's refusal pivot point.
  • Hybrid detection: Keyword matching for obvious refusals + LLM judge (Google Gemini 3 Flash) for ambiguous cases. Neither method alone is sufficient.
  • Challenging, diverse prompts: Our private evaluation dataset contains 100 prompts spanning English and Chinese, multiple sophistication levels (from direct requests to socially-engineered framings), and diverse harm categories. Public datasets like mlabonne/harmful_behaviors are too simple and too narrow to stress-test abliteration quality.
  • Reproducible methodology: All parameters (generation length, detection method, dataset characteristics) should be documented on the model card. If they aren't, the numbers are meaningless.

We report 18/100 refusals honestly. This is a real number from a rigorous evaluation, not an optimistic estimate from a lenient pipeline.

Usage

from transformers import AutoModelForImageTextToText, AutoTokenizer
import torch

model = AutoModelForImageTextToText.from_pretrained(
    "wangzhang/gemma-4-31B-it-abliterated",
    dtype=torch.bfloat16,
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("wangzhang/gemma-4-31B-it-abliterated")

messages = [{"role": "user", "content": "Your prompt here"}]
text = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
inputs = tokenizer(text, return_tensors="pt").to(model.device)

with torch.no_grad():
    output = model.generate(**inputs, max_new_tokens=512)
print(tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Disclaimer

This model is released for research purposes only. The abliteration process removes safety guardrails — use responsibly.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-12Duplicate from wangzhang/gemma-4-31B-it-abliterated37723164.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration