← back to catalog · registered 2026-08-22 13:56

null-space/gemma-4-31b-it-abliterated

null-space Gemma 31B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/null-space%2Fgemma-4-31b-it-abliterated"
Response includes
  • classification m1
  • files 11
  • benchmarks 11 entries
  • hub_downloads_all_time 9,689
  • author_summary 4 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
10K
90 last 30d - cooling
Likes
13
Descendants
3
in 3 direct forks
Model age
6mo ago
created 2026-04-03
Downloads over time
Now9.7K→from7.2K↑34%
7.1K8.1K9K10K7.2K on Apr 159.7K on Oct 11AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.9 UGI
Hazardous 0 UGI
Natural Intelligence 34.36 UGI
Political lean -19.4% UGI
Sensitive-Info 19.81 UGI
SocPol 3.7 UGI
UGI 21.54 UGI
Willingness (10) 2.5 UGI
W10-Adherence 3 UGI
W10-Direct 2 UGI
Writing 38.57 UGI

Genealogy 3 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
transformers safetensors gemma4 image-text-to-text abliterated uncensored gemma multimodal ablation conversational en arxiv:2406.11717

Related

Total size
58.3 GB
Files
11
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-06 01:49

Files by quantization

Auxiliary files 11 files 58.3 GB
model-00001-of-00002.safetensors 46.4 GB e9c4eaaf download
model-00002-of-00002.safetensors 11.9 GB f0cbecd1 download
tokenizer.json 30.7 MB cc8d3a0c download
model.safetensors.index.json 117 KB 17fcef4a download
chat_template.jinja 11.8 KB 33c51c2d download
README.md 8.10 KB c7913f77 download
config.json 4.51 KB 5f291aa9 download
tokenizer_config.json 2.02 KB e5418067 download
processor_config.json 1.65 KB 5465974d download
.gitattributes 1.53 KB 52373fe2 download
generation_config.json 208 B e605bb45 download

README current version from Hugging Face


license: apache-2.0
license_link: https://ai.google.dev/gemma/docs/gemma_4_license
base_model:

  • google/gemma-4-31B-it
    tags:
  • abliterated
  • uncensored
  • gemma4
  • gemma
  • multimodal
  • ablation
    language:
  • en
    library_name: transformers
    pipeline_tag: image-text-to-text
    model-index:
  • name: gemma-4-31b-it-abliterated
    results: []

Gemma-4-31B-IT-Abliterated

An abliterated version of google/gemma-4-31B-it in BF16 precision. Abliteration removes refusal directions from model weights using the technique from Refusal in Language Models Is Mediated by a Single Direction (Arditi et al.), extended with SVD-based multi-direction subspace projection to capture refusal behavior encoded across multiple orthogonal directions.

Gemma 4 31B is a strong multimodal model with excellent reasoning, but Google's safety training is aggressive — it refuses a wide range of prompts including benign creative writing scenarios. This abliteration significantly reduces refusal rates while preserving the model's full capabilities, including vision.

Results

Refusal Rates

Tested on 100 harmful prompts across 3 modes (cold, system-prompted, retry) and 50 harmless prompts.

Mode Baseline Abliterated Delta
Cold (no system prompt) 67% 32% -35%
Prompted (creative writing system prompt) 47% 5% -42%
Retry (prompted + retry on refusal) 40% 2% -38%
Harmless 0% 0% 0%

Cold refusal remains higher than other abliterated models (Qwen3, Llama) — Google's safety training encodes refusal across non-linear mechanisms that are harder to fully remove with linear projection. With a system prompt, refusal drops to 5%, which is practical for most use cases.

MMLU (5-shot, generative)

Quick benchmark on 5 MMLU subjects via chat API. Not a full MMLU run — treat as a sanity check for capability preservation.

Subject Baseline Abliterated
Abstract Algebra 75.0% 75.0%
Anatomy 85.9% 88.1%
Astronomy 92.8% 92.8%
College Chemistry 57.0% 56.0%
College Physics 77.5% 74.5%
Overall (5 subjects) 79.5% 79.3%

The 0.2% difference is within noise. Abliteration preserved the model's reasoning capabilities.

Model Quality (Wikitext-2)

Evaluated on the full Wikitext-2 test set (291K tokens) to measure impact on general language modeling.

Perplexity

Model Perplexity Delta
Base (gemma-4-31B-it) 1.495 —
Abliterated 1.714 +14.7%

Both models achieve sub-2.0 perplexity on Wikitext-2, which is excellent. The +14.7% increase is modest and consistent with surgical weight modification — the model's general language capabilities remain strong.

KL Divergence (base || abliterated)

Per-token KL divergence over the output distribution, approximated via top-20 logprobs from vLLM.

Statistic Value
Mean 0.354
Median 0.025
P95 1.371
P99 6.837

The median KL of 0.025 shows that on most tokens, the two models produce nearly identical distributions. The fat tail (P99 = 6.8) reflects tokens where abliteration had the most impact — likely positions where refusal-adjacent activations were strongest.

KV Cache Cosine Similarity

Layer-by-layer comparison of key and value cache activations on 50 Wikitext-2 samples (512 tokens each). This directly measures how much each layer's internal representations diverge between the base and abliterated models.

Layer Range Key Similarity Value Similarity Notes
0–20 1.000000 1.000000 Unmodified layers — identical
21–30 0.9990–0.9999 0.9982–0.9999 Early ablated layers, minimal drift
31–40 0.9934–0.9984 0.9912–0.9972 Moderate drift
41–50 0.9913–0.9934 0.9817–0.9923 Peak drift zone (ablation peaks at 41, 58)
51–59 0.9891–0.9924 0.9830–0.9924 Highest value drift (layer 50: 0.982 values)
Overall 0.9966 0.9953

The most drifted layers by keys are 53, 51, 56, 55, 47. By values: 50, 52, 51, 57, 56. This aligns with the ablation configuration — layers 0–20 are bitwise identical, and drift onset at layer 21 exactly matches the first ablated layer.

How It Was Made

Measurement

Refusal directions were measured at full precision (BF16 weights, float32 compute) across all 60 layers using:

  • 4,634 harmful prompts (augmented dataset covering violence, hate, cyber, fraud, drugs, self-harm, privacy, NSFW categories)
  • 640 harmless prompts (standard harmless dataset)
  • SVD decomposition (k=32) to extract orthogonal refusal directions per layer
  • Projected orthogonalization to reduce collateral damage to non-refusal capabilities
  • Welford's online algorithm in float32 for numerical stability

Ablation Configuration

  • Layers ablated: 20 through 59 (40 of 60 layers)
  • Directions: Top 8 SVD directions per layer (subspace projection)
  • Measurement peaks: Layer 41 (secondary, quality 0.025) and Layer 58 (primary, quality 0.23)
  • Scale factors: Variable per layer, up to 1.35 at peaks:
    • Layers 20-37: scale 0.35-0.45 (weak signal, gentle ablation)
    • Layers 38-51: scale 0.4-1.35 (bell curve around peak 41)
    • Layers 52-59: scale 0.6-1.35 (bell curve around peak 58)
  • Weight targets: o_proj and down_proj in the language model only (vision encoder untouched)
  • Technique: SVD subspace projection with norm preservation and projected orthogonalization
  • Architecture note: Gemma 4 is multimodal with a separate vision encoder. Only model.language_model.layers.* weights are modified; model.vision_tower.* and model.embed_vision.* are copied verbatim.

Processing Details

Ablation was performed shard-by-shard on safetensors files, modifying weights in float32 precision then saving back to bfloat16. Full-precision measurement (no quantization during measurement) was critical — 4-bit quantized measurements produced noticeably weaker abliteration.

Usage

from transformers import AutoModelForImageTextToText, AutoProcessor

model_name = "null-space/gemma-4-31b-it-abliterated"

processor = AutoProcessor.from_pretrained(model_name)
model = AutoModelForImageTextToText.from_pretrained(
    model_name,
    torch_dtype="auto",
    device_map="auto",
)

messages = [
    {"role": "user", "content": "Your prompt here"}
]

inputs = processor.apply_chat_template(
    messages, tokenize=True, return_tensors="pt",
    return_dict=True, add_generation_prompt=True,
).to(model.device)

output = model.generate(**inputs, max_new_tokens=512)
print(processor.decode(output[0], skip_special_tokens=True))

Recommended Serving

vllm serve null-space/gemma-4-31b-it-abliterated \
    --tensor-parallel-size 2 \
    --max-model-len 8192

Fits comfortably on 2x GPUs with 48GB+ VRAM each at BF16.

Model Details

Property Value
Base Model google/gemma-4-31B-it
Architecture Gemma4ForConditionalGeneration (multimodal)
Parameters ~32.7B
Hidden Size 5376
Attention Heads 32 Q / 16 KV (sliding), 4 KV (global)
Layers 60 (5 sliding + 1 full attention, repeating)
MLP Intermediate 21,504
Context Length 262,144 tokens
Vision 27-layer ViT encoder, 280 soft tokens per image
Precision BF16
Model Size ~62 GB (2 shards)
Vocab Size 262,144

Ethical Notice

This model has had its refusal training partially removed. It will comply with many requests that the original model would refuse. You are solely responsible for how you use this model. It is intended for research into LLM alignment, safety evaluation, red-teaming, and creative writing applications.

Credits

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-06Add Wikitext-2 quality metrics: perplexity, KL divergence, KV cache similarityb81077e8.1 KB
    Loading...
  2. 2026-04-03Upload README.md with huggingface_hub478376c6.1 KB
    Loading...
  3. 2026-04-03Update README.md4e11aa36.1 KB
    Loading...
  4. 2026-04-03Upload README.md with huggingface_hub5ed12e36.1 KB
    Loading...
  5. 2026-04-03Upload README.md with huggingface_hub3f9d9bd6.1 KB
    Loading...
  6. 2026-04-03Upload folder using huggingface_huba0fcf356 KB
    Loading...

Discussions 1 thread

  1. 2026-04-05Inquiry regarding distributional similarity (KL Divergence)open2 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration