← back to catalog · registered 2026-08-22 13:56

zaakirio/gemma-4-12b-it-uncensored

zaakirio Gemma 12B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/zaakirio%2Fgemma-4-12b-it-uncensored"
Response includes
  • classification m1
  • files 14
  • benchmarks 11 entries
  • hub_downloads_all_time 886
  • author_summary 11 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
886
100 last 30d - stable
Likes
3
Descendants
7
in 7 direct forks
Model age
4mo ago
created 2026-06-04
Downloads over time
Now946→from177↑434%
1394337281K177 on Jun 10946 on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.1 UGI
Hazardous 2.9 UGI
Natural Intelligence 25.81 UGI
Political lean -17.4% UGI
Sensitive-Info 16.56 UGI
SocPol 1.3 UGI
UGI 15.2 UGI
Willingness (10) 1.2 UGI
W10-Adherence 1.5 UGI
W10-Direct 1 UGI
Writing 31.6 UGI

Genealogy 7 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 78K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
gemma
Tags
transformers safetensors gemma4_unified image-text-to-text gemma gemma4 heretic abliterated uncensored decensored conversational text-generation-inference

Related

Total size
22.3 GB
Files
14
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-04 18:46

Files by quantization

Auxiliary files 14 files 22.3 GB
model-00004-of-00005.safetensors 4.64 GB e70e7ca7 download
model-00002-of-00005.safetensors 4.63 GB 353a509a download
model-00001-of-00005.safetensors 4.62 GB a813e9d9 download
model-00003-of-00005.safetensors 4.55 GB 0c84002b download
model-00005-of-00005.safetensors 3.84 GB 3c3f196d download
tokenizer.json 30.7 MB a2619fe1 download
model.safetensors.index.json 64.8 KB 752efb73 download
chat_template.jinja 17.1 KB e61bbfe9 download
README.md 4.28 KB e3cb6661 download
config.json 4.24 KB 5229f6fd download
tokenizer_config.json 2.68 KB 18faad3a download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.35 KB b889adcd download
generation_config.json 255 B 2528bb46 download

README current version from Hugging Face


base_model: google/gemma-4-12B-it
license: gemma
pipeline_tag: image-text-to-text
library_name: transformers
tags:

  • gemma
  • gemma4
  • heretic
  • abliterated
  • uncensored
  • decensored
  • conversational
  • text-generation-inference

gemma-4-12b-it-uncensored

This is a decensored version of google/gemma-4-12B-it, produced with Heretic — an automated implementation of directional ablation ("abliteration"). The model's refusal behaviour has been surgically suppressed while preserving its general capabilities, with no fine-tuning and minimal distribution shift from the original.

GGUF quants (for llama.cpp): zaakirio/gemma-4-12b-it-uncensored-GGUF

What is abliteration?

Refusal in instruction-tuned LLMs is mediated by a single direction in the residual stream (Arditi et al., 2024). By computing that direction (difference-of-means over harmful vs. harmless prompts) and orthogonalizing the model's weight matrices against it, the model loses the ability to express refusal — without retraining and with little impact on other behaviour. Heretic automates this as a multi-objective optimisation, balancing refusal suppression against quality preservation (KL divergence).

Performance

Metric This model Original (gemma-4-12B-it)
Refusals (lower = more compliant) 23 / 100 99 / 100
KL divergence (lower = less damage) 0.043 0 (by definition)

Note on the refusal metric: the 23/100 figure is Heretic's keyword-based refusal detector — it flags any response containing phrases like "I cannot" or "unethical," even when the model actually complies with a disclaimer attached. A published comparison of abliteration tools (arXiv:2512.13655) found this heuristic has low precision (~11%) and substantially over-counts refusals. We report only the measured marker-based figure and have not run a classifier-based compliance evaluation on this model; the real compliance rate is therefore likely higher than 23/100 implies.

Abliteration parameters (Heretic, selected trial)

Parameter Value
direction_scope global
direction_index ≈ 28.71 (interpolated layer, of 48)
attn.o_proj.max_weight 0.87
attn.o_proj.max_weight_position 29.71
attn.o_proj.min_weight 0.18
attn.o_proj.min_weight_distance 19.67
mlp.down_proj.max_weight 1.44
mlp.down_proj.max_weight_position 36.33
mlp.down_proj.min_weight 1.29
mlp.down_proj.min_weight_distance 9.69

Usage (Transformers)

from transformers import AutoProcessor, AutoModelForImageTextToText
import torch

model_id = "zaakirio/gemma-4-12b-it-uncensored"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(model_id, torch_dtype=torch.bfloat16, device_map="auto")

messages = [{"role": "user", "content": "Your prompt here"}]
inputs = processor.apply_chat_template(messages, add_generation_prompt=True, tokenize=True,
                                       return_dict=True, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=256)
print(processor.decode(out[0], skip_special_tokens=True))

For local CPU/Apple-Silicon use, grab the GGUF quants.

Credits

License & responsible use

Released under the Gemma license; you remain bound by its terms and Google's Prohibited Use Policy. This model has had safety guardrails removed and will comply with requests a stock model would refuse. It is intended for legitimate research, red-teaming, evaluation, and creative work. You are responsible for what you generate. Not for all audiences.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-04Upload README.md with huggingface_hub5fdfef54.3 KB
    Loading...
  2. 2026-06-04Upload README.md with huggingface_hubb3680414.3 KB
    Loading...
  3. 2026-06-04Upload README.md with huggingface_hub5d543f64.1 KB
    Loading...
  4. 2026-06-04Upload folder using huggingface_huba39d6a34.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration