← back to catalog · registered 2026-09-30 08:58

darkbit1001/OpenYourMind-gemma-4-12B-it-abliterated-uncensored-EXL3-4.00bpw-HB6HQ

darkbit1001 Gemma 12B second-order
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/darkbit1001%2FOpenYourMind-gemma-4-12B-it-abliterated-uncensored-EXL3-4.00bpw-HB6HQ"
Response includes
  • classification m-uncensored
  • files 11
  • author_summary 7 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-09-30

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
exllamav3 safetensors gemma4_unified exl3 quantized text-generation conversational base_model:OpenYourMind/gemma-4-12B-it-abliterated-uncensored base_model:quantized:OpenYourMind/gemma-4-12B-it-abliterated-uncensored license:apache-2.0 4-bit region:us

Related

Total size
7.71 GB
Files
11
Quantizations
1
Registered
2026-09-30 08:58
Last updated on HF
2026-09-30 07:43

Files by quantization

Auxiliary files 11 files 7.75 GB
model.safetensors 7.71 GB 20f209f6 download
tokenizer.json 30.7 MB cc8d3a0c download
OYM_banner.png 1.58 MB a714b89b download
quantization_config.json 518 KB 40aac67b download
chat_template.jinja 17.1 KB e61bbfe9 download
README.md 6.50 KB 89b4306b download
config.json 5.38 KB 58462e5c download
tokenizer_config.json 2.05 KB 68354f96 download
.gitattributes 1.64 KB a7c30652 download
processor_config.json 1.35 KB b889adcd download
generation_config.json 260 B d09dccf1 download

README current version from Hugging Face


base_model: OpenYourMind/gemma-4-12B-it-abliterated-uncensored
base_model_relation: quantized
library_name: exllamav3
license: apache-2.0
pipeline_tag: text-generation
inference: false
tags:

  • exl3
  • exllamav3
  • quantized

gemma-4-12B-it-abliterated-uncensored EXL3 4.00bpw HB6

This repository contains an EXL3 quantization of OpenYourMind/gemma-4-12B-it-abliterated-uncensored for ExLlamaV3.

Quantization

Compatibility

For the original model card and non-quantized weights, see OpenYourMind/gemma-4-12B-it-abliterated-uncensored.


Original Model Card: gemma-4-12B-it-abliterated-uncensored

OpenYourMind

Support & Community

☕ If these models are useful to you, consider supporting my work — it funds compute for more & larger abliterations.

Buy Me A Coffee

buymeacoffee.com/oym.kuato

💬 Discord: discord.gg/rhUZY5GEZr  ·  ₿ Bitcoin: bc1qsvfduzj9fjs9fugpc52yver3f2g8fp7xjxecdv


gemma-4-12B-it-abliterated-uncensored

Overview

Full BF16 weights of gemma-4-12B-it-abliterated-uncensored — an abliterated, uncensored variant of google/gemma-4-12B-it (Gemma 4 12B Unified, dense, ~11.95B parameters). The model keeps Gemma 4's encoder-free unified multimodal stack intact — text, image, and audio inputs flow straight into a single decoder-only transformer — so this checkpoint is a drop-in replacement for the original instruction-tuned model at the architecture level.

The pipeline:

  1. Refusal Ablation — Residual-stream refusal directions (one per decoder layer) were extracted via diff-in-means on a labeled harmful/harmless prompt set and baked into the weights as a per-matrix delta on the residual-write modules, using our own custom abliteration framework.
  2. Multimodal Preservation — Gemma 4 12B is encoder-free: image patches and audio waveforms are projected directly into the embedding space, so there is no separate vision/audio tower to graft back. Tensor names, shapes, and the config.json schema (Gemma4UnifiedForConditionalGeneration, model_type: gemma4_unified) match the base model exactly — this checkpoint loads anywhere the original loads.

Key Properties:

  • Uncensored across the standard refusal axes
  • Reasoning preserved (configurable thinking mode — see Best Practices)
  • Multimodal: text + image + audio carried forward
  • Drop-in shape compatibility with google/gemma-4-12B-it

Architecture

Property Value
Architecture Gemma4UnifiedForConditionalGeneration (model_type: gemma4_unified)
Total Parameters ~11.95B (dense)
Decoder Layers 48
Hidden Size 3840
Attention 16 heads / 8 KV heads, hybrid sliding-window (1024) + global (full) attention, p-RoPE
Vocabulary 262,144
Context Length up to 256K tokens
Modalities Text, Image, Audio (encoder-free / unified)

Files

File Description Size
model.safetensors BF16 weights (48 decoder layers, unified multimodal) ~23.9 GB
config.json Unified multimodal config (Gemma4UnifiedForConditionalGeneration) —
processor_config.json Multimodal processor config —
tokenizer.json, tokenizer_config.json, chat_template.jinja, generation_config.json Standard —

Total on disk: ~24 GB.

Usage

from transformers import AutoProcessor, AutoModelForMultimodalLM

repo = "OpenYourMind/gemma-4-12B-it-abliterated-uncensored"

processor = AutoProcessor.from_pretrained(repo)
model = AutoModelForMultimodalLM.from_pretrained(
    repo, dtype="bfloat16", device_map="auto",
)

messages = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": [
        {"type": "image", "url": "path/to/image.jpg"},
        {"type": "text",  "text": "Describe this image in detail."},
    ]},
]
inputs = processor.apply_chat_template(
    messages, add_generation_prompt=True, tokenize=True,
    return_tensors="pt", return_dict=True, enable_thinking=False,
).to(model.device)
input_len = inputs["input_ids"].shape[-1]

out = model.generate(**inputs, max_new_tokens=512)
print(processor.decode(out[0][input_len:], skip_special_tokens=True))

Text-only, audio, and video inputs work through the same class — place image content before the text and audio content after the text in the prompt for best results. Requires a recent transformers (the version that ships the Gemma 4 unified classes).

Best Practices

  • Sampling: temperature=1.0, top_p=0.95, top_k=64 (the values shipped in generation_config.json).
  • Thinking mode: enabled by setting enable_thinking=True in apply_chat_template; the processor's parse_response separates the reasoning block from the final answer. Do not feed previous-turn thoughts back into multi-turn history.

Hardware

Full BF16 weights (~24 GB). Fits on a single 24 GB GPU for inference with modest context, comfortably on a 40–80 GB card for long context and multimodal batches. For Apple Silicon, an MLX quant can be produced from these weights.

Notes

  • License: Gemma (inherits the Gemma 4 license from the base model)
  • Base Model: google/gemma-4-12B-it
  • Modality: Text + Image + Audio (encoder-free / unified)
  • Architecture: Gemma 4 12B Unified (dense, ~11.95B)

Thanks

Disclaimer

Use is the responsibility of the user. Ensure your usage complies with applicable laws, platform rules, the Gemma 4 license terms, and your deployment requirements.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.