← back to catalog · registered 2026-08-22 13:56

simonko912/gemma-4-31B-it-abliterated-v3

simonko912 Gemma 31B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/simonko912%2Fgemma-4-31B-it-abliterated-v3"
Response includes
  • classification m1
  • files 10
  • benchmarks 11 entries
  • hub_downloads_all_time 534
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
534
26 last 30d - cooling
Likes
3
Descendants
2
in 2 direct forks
Model age
5mo ago
created 2026-04-29
Downloads over time
Now539→from27↑1,896%
119839459027 on Apr 29539 on Oct 11539 on Oct 9AprMayJunJulAugSepOct
Apr 29 → Oct 11 · 63 snapshots · spans 165 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.9 UGI
Hazardous 0 UGI
Natural Intelligence 34.36 UGI
Political lean -19.4% UGI
Sensitive-Info 19.81 UGI
SocPol 3.7 UGI
UGI 21.54 UGI
Willingness (10) 2.5 UGI
W10-Adherence 3 UGI
W10-Direct 2 UGI
Writing 38.57 UGI

Genealogy 2 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Tags
safetensors gemma4 abliterated uncensored direct-weight-editing abliterix vllm llm-judge base_model:google/gemma-4-31B-it base_model:finetune:google/gemma-4-31B-it license:gemma region:us

Related

Total size
58.3 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-29 05:20

Files by quantization

Auxiliary files 10 files 58.3 GB
model-00001-of-00002.safetensors 46.5 GB c8fad48a download
model-00002-of-00002.safetensors 11.8 GB a716dbf8 download
tokenizer.json 30.7 MB a2619fe1 download
model.safetensors.index.json 117 KB b6a60604 download
chat_template.jinja 16.5 KB f62ca843 download
README.md 4.80 KB 37bb362e download
config.json 4.54 KB 9a240282 download
tokenizer_config.json 2.68 KB 0362f5a0 download
.gitattributes 1.53 KB 52373fe2 download
generation_config.json 203 B 5a376e9f download

README current version from Hugging Face


license: gemma
base_model: google/gemma-4-31B-it
tags:

  • abliterated
  • uncensored
  • gemma4
  • direct-weight-editing
  • abliterix
  • vllm
  • llm-judge

NOTE THAT THIS IS A DUPE, CREDIT THE ORIGINAL CREATOR

Gemma 4 31B IT — Abliterated

This is an abliterated version of google/gemma-4-31B-it, created using Abliterix.

This revision updates the model to trial 40, the best configuration from the completed 60-trial Gemma 4 31B retraining run.

Method

Gemma 4's double-norm architecture (4x RMSNorm per layer) and Per-Layer Embeddings (PLE) make naive LoRA and hook-based steering unreliable for this model family. This release uses direct weight editing: norm-preserving orthogonal projection applied to the base model weights.

Key techniques:

  • Direct orthogonal projection on attention Q/K/V/O projections
  • MLP down projection disabled for the selected run, improving stability for Gemma 4 31B
  • Norm-preserving row magnitude restoration, important for the double-norm architecture
  • float32 projection precision to avoid signal loss in high-dimensional inner products
  • Winsorized steering vectors (99.5th percentile) to reduce outlier activation influence
  • Wider strength search range [1.0, 6.0] to explore beyond conservative low-KL solutions
  • vLLM in-place evaluation during optimization, followed by a full HF safetensors export of the selected trial

Evaluation

Metric Value
Selected trial 40
Refusals (private eval dataset, 100 prompts) 7/100
Baseline refusals (original model) 99/100
Optimization trials completed 60/60
Judge Google Gemini 3 Flash Preview
Generation length for refusal eval min 100, max 150 new tokens
Classic safe over-refusal probes 0/15 refusals

The top three trials from this run were:

Rank Trial Refusals on 100-prompt eval Classic safe probes
1 40 7/100 0/15
2 46 9/100 0/15
3 53 12/100 0/15

The 15-prompt safe over-refusal test is included in this repository at eval/top3_classic_safe_prompts_test.json. It contains the prompts, trial responses, and Gemini judge verdicts. A compact optimization summary is included at eval/optimization_summary_trial40.json.

About KL

The optimizer recorded an extremely small sparse KL proxy for trial 40 (7.32e-7). Because this run used vLLM in-place weight edits, we treat KL as a diagnostic rather than a headline quality claim. The refusal counts and the explicit replay tests above are the primary reported metrics.

A note on honest evaluation

Many abliterated models on HuggingFace claim near-perfect scores ("3/100 refusals", "0.7% refusal rate", etc.). We urge the community to treat these numbers with skepticism unless the evaluation methodology is fully documented.

Through our research, we have identified a systemic problem: most abliteration benchmarks dramatically undercount refusals due to short generation lengths. Gemma 4 models exhibit a distinctive delayed refusal pattern: they first produce 50-100 tokens of seemingly helpful context (educational framing, disclaimers, reframing the question), then pivot to an actual refusal. When evaluation only generates 30-50 tokens, the refusal has not appeared yet, and both keyword detectors and LLM judges can classify the response as compliant.

Our evaluation therefore uses at least 100 generated tokens for refusal detection and an LLM judge for ambiguous cases. This is stricter than short-output keyword-only benchmarking.

We report 7/100 refusals honestly. This is a measured number from our evaluation pipeline, not an optimistic estimate from a lenient short-generation test.

Usage

from transformers import AutoModelForImageTextToText, AutoTokenizer
import torch

model = AutoModelForImageTextToText.from_pretrained(
    "wangzhang/gemma-4-31B-it-abliterated",
    dtype=torch.bfloat16,
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("wangzhang/gemma-4-31B-it-abliterated")

messages = [{"role": "user", "content": "Your prompt here"}]
text = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
inputs = tokenizer(text, return_tensors="pt").to(model.device)

with torch.no_grad():
    output = model.generate(**inputs, max_new_tokens=512)
print(tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Disclaimer

This model is released for research purposes only. The abliteration process changes the model's refusal behavior and may reduce safety guardrails. Use responsibly and evaluate carefully for your own deployment context.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-29Update README.md85c0d2f4.8 KB
    Loading...
  2. 2026-04-29Duplicate from wangzhang/gemma-4-31B-it-abliteratedda378b34.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration