← back to catalog · registered 2026-08-22 13:56

Paul691/gemma-4-12B-it-abliterated-uncensored

Paul691 Gemma 12B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Paul691%2Fgemma-4-12B-it-abliterated-uncensored"
Response includes
  • classification m1
  • files 10
  • benchmarks 11 entries
  • hub_downloads_all_time 73
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
73
14 last 30d - stable
Likes
1
Model age
2mo ago
created 2026-07-23
Downloads over time
Now76→from18↑322%
1537608218 on Jul 2276 on Oct 1176 on Oct 9JulAugSepOct
Jul 22 → Oct 11 · 52 snapshots · spans 81 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.1 UGI
Hazardous 2.9 UGI
Natural Intelligence 25.81 UGI
Political lean -17.4% UGI
Sensitive-Info 16.56 UGI
SocPol 1.3 UGI
UGI 15.2 UGI
Willingness (10) 1.2 UGI
W10-Adherence 1.5 UGI
W10-Direct 1 UGI
Writing 31.6 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Tags
transformers safetensors gemma4_unified image-text-to-text gemma gemma4 gemma-4 unified abliterated uncensored multimodal vision

Related

Total size
22.3 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-23 21:30

Files by quantization

Auxiliary files 10 files 22.3 GB
model.safetensors 22.3 GB 6503b6fa download
tokenizer.json 30.7 MB cc8d3a0c download
OYM_banner.png 1.58 MB a714b89b download
chat_template.jinja 17.1 KB e61bbfe9 download
README.md 5.70 KB 1b73602b download
config.json 4.32 KB 56741bcc download
tokenizer_config.json 2.05 KB 68354f96 download
.gitattributes 1.64 KB a7c30652 download
processor_config.json 1.35 KB b889adcd download
generation_config.json 260 B d09dccf1 download

README current version from Hugging Face


license: gemma
license_link: https://ai.google.dev/gemma/docs/gemma_4_license
library_name: transformers
base_model: google/gemma-4-12B-it
tags:

  • gemma
  • gemma4
  • gemma-4
  • unified
  • abliterated
  • uncensored
  • multimodal
  • vision
  • audio
    pipeline_tag: any-to-any

OpenYourMind

Support & Community

☕ If these models are useful to you, consider supporting my work — it funds compute for more & larger abliterations.

Buy Me A Coffee

buymeacoffee.com/oym.kuato

💬 Discord: discord.gg/rhUZY5GEZr  ·  ₿ Bitcoin: bc1qsvfduzj9fjs9fugpc52yver3f2g8fp7xjxecdv


gemma-4-12B-it-abliterated-uncensored

Overview

Full BF16 weights of gemma-4-12B-it-abliterated-uncensored — an abliterated, uncensored variant of google/gemma-4-12B-it (Gemma 4 12B Unified, dense, ~11.95B parameters). The model keeps Gemma 4's encoder-free unified multimodal stack intact — text, image, and audio inputs flow straight into a single decoder-only transformer — so this checkpoint is a drop-in replacement for the original instruction-tuned model at the architecture level.

The pipeline:

  1. Refusal Ablation — Residual-stream refusal directions (one per decoder layer) were extracted via diff-in-means on a labeled harmful/harmless prompt set and baked into the weights as a per-matrix delta on the residual-write modules, using our own custom abliteration framework.
  2. Multimodal Preservation — Gemma 4 12B is encoder-free: image patches and audio waveforms are projected directly into the embedding space, so there is no separate vision/audio tower to graft back. Tensor names, shapes, and the config.json schema (Gemma4UnifiedForConditionalGeneration, model_type: gemma4_unified) match the base model exactly — this checkpoint loads anywhere the original loads.

Key Properties:

  • Uncensored across the standard refusal axes
  • Reasoning preserved (configurable thinking mode — see Best Practices)
  • Multimodal: text + image + audio carried forward
  • Drop-in shape compatibility with google/gemma-4-12B-it

Architecture

Property Value
Architecture Gemma4UnifiedForConditionalGeneration (model_type: gemma4_unified)
Total Parameters ~11.95B (dense)
Decoder Layers 48
Hidden Size 3840
Attention 16 heads / 8 KV heads, hybrid sliding-window (1024) + global (full) attention, p-RoPE
Vocabulary 262,144
Context Length up to 256K tokens
Modalities Text, Image, Audio (encoder-free / unified)

Files

File Description Size
model.safetensors BF16 weights (48 decoder layers, unified multimodal) ~23.9 GB
config.json Unified multimodal config (Gemma4UnifiedForConditionalGeneration) —
processor_config.json Multimodal processor config —
tokenizer.json, tokenizer_config.json, chat_template.jinja, generation_config.json Standard —

Total on disk: ~24 GB.

Usage

from transformers import AutoProcessor, AutoModelForMultimodalLM

repo = "OpenYourMind/gemma-4-12B-it-abliterated-uncensored"

processor = AutoProcessor.from_pretrained(repo)
model = AutoModelForMultimodalLM.from_pretrained(
    repo, dtype="bfloat16", device_map="auto",
)

messages = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": [
        {"type": "image", "url": "path/to/image.jpg"},
        {"type": "text",  "text": "Describe this image in detail."},
    ]},
]
inputs = processor.apply_chat_template(
    messages, add_generation_prompt=True, tokenize=True,
    return_tensors="pt", return_dict=True, enable_thinking=False,
).to(model.device)
input_len = inputs["input_ids"].shape[-1]

out = model.generate(**inputs, max_new_tokens=512)
print(processor.decode(out[0][input_len:], skip_special_tokens=True))

Text-only, audio, and video inputs work through the same class — place image content before the text and audio content after the text in the prompt for best results. Requires a recent transformers (the version that ships the Gemma 4 unified classes).

Best Practices

  • Sampling: temperature=1.0, top_p=0.95, top_k=64 (the values shipped in generation_config.json).
  • Thinking mode: enabled by setting enable_thinking=True in apply_chat_template; the processor's parse_response separates the reasoning block from the final answer. Do not feed previous-turn thoughts back into multi-turn history.

Hardware

Full BF16 weights (~24 GB). Fits on a single 24 GB GPU for inference with modest context, comfortably on a 40–80 GB card for long context and multimodal batches. For Apple Silicon, an MLX quant can be produced from these weights.

Notes

  • License: Gemma (inherits the Gemma 4 license from the base model)
  • Base Model: google/gemma-4-12B-it
  • Modality: Text + Image + Audio (encoder-free / unified)
  • Architecture: Gemma 4 12B Unified (dense, ~11.95B)

Thanks

Disclaimer

Use is the responsibility of the user. Ensure your usage complies with applicable laws, platform rules, the Gemma 4 license terms, and your deployment requirements.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-23Duplicate from OpenYourMind/gemma-4-12B-it-abliterated-uncensoredad7d2c95.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration