license: apache-2.0
base_model:
- google/gemma-4-12B-it
pipeline_tag: image-text-to-text
tags: - ultragemma4
- text-generation-inference
- llama-cpp
- decensored
- abliterated
- unfiltered
- unredacted
- heretic
language: - en
library_name: transformers
Use Q4_K_S or higher for standard performance. Q4_K_M is recommended.
Key Highlights
- Heretic-Based Abliteration: Modified using the Heretic toolkit to identify and alter refusal-related representations within the model.
- Reduced Refusal Behavior: Optimized to minimize internal refusal tendencies while maintaining instruction-following capabilities.
- Gemma 4 12B Unified Backbone: Built directly on top of google/gemma-4-12B-it.
- Multimodal Foundation: Inherits native text, image, audio, and video understanding capabilities from the Gemma 4 Unified architecture.
- Reasoning-Oriented Performance: Preserves multi-step reasoning and analytical capabilities after abliteration.
- Research-Focused Release: Designed for alignment research, model behavior analysis, and evaluation of refusal-direction modifications.
- 12B Scale Deployment: Suitable for local inference, research environments, and optimized deployment setups.
Model Files
| File Name | Quant Type | File Size | File Link |
|---|---|---|---|
| ultragemma4-12b-heretic-uncensored.BF16.gguf | BF16 | 23.8 GB | Download |
| ultragemma4-12b-heretic-uncensored.F16.gguf | F16 | 23.8 GB | Download |
| ultragemma4-12b-heretic-uncensored.Q2_K.gguf | Q2_K | 4.83 GB | Download |
| ultragemma4-12b-heretic-uncensored.Q3_K_L.gguf | Q3_K_L | 6.57 GB | Download |
| ultragemma4-12b-heretic-uncensored.Q3_K_M.gguf | Q3_K_M | 6.09 GB | Download |
| ultragemma4-12b-heretic-uncensored.Q3_K_S.gguf | Q3_K_S | 5.53 GB | Download |
| ultragemma4-12b-heretic-uncensored.Q4_0.gguf | Q4_0 | 6.98 GB | Download |
| ultragemma4-12b-heretic-uncensored.Q4_K_M.gguf | Q4_K_M | 7.38 GB | Download |
| ultragemma4-12b-heretic-uncensored.Q4_K_S.gguf | Q4_K_S | 7.02 GB | Download |
| ultragemma4-12b-heretic-uncensored.Q5_0.gguf | Q5_0 | 8.34 GB | Download |
| ultragemma4-12b-heretic-uncensored.Q5_K_M.gguf | Q5_K_M | 8.55 GB | Download |
| ultragemma4-12b-heretic-uncensored.Q5_K_S.gguf | Q5_K_S | 8.34 GB | Download |
| ultragemma4-12b-heretic-uncensored.Q6_K.gguf | Q6_K | 9.79 GB | Download |
| ultragemma4-12b-heretic-uncensored.Q8_0.gguf | Q8_0 | 12.7 GB | Download |
| ultragemma4-12b-heretic-uncensored.mmproj-bf16.gguf | mmproj-bf16 | 175 MB | Download |
| ultragemma4-12b-heretic-uncensored.mmproj-f16.gguf | mmproj-f16 | 175 MB | Download |
| ultragemma4-12b-heretic-uncensored.mmproj-q8_0.gguf | mmproj-q8_0 | 159 MB | Download |
Quick Start with llama.cpp (Docker)
FROM ghcr.io/ggml-org/llama.cpp:full
WORKDIR /app
RUN apt update && apt install -y python3-pip
RUN pip install -U huggingface_hub --break-system-packages
RUN python3 -c 'from huggingface_hub import hf_hub_download; \
repo="prithivMLmods/gemma-4-12B-it-heretic_decensored-GGUF"; \
hf_hub_download(repo_id=repo, filename="gemma-4-12B-it-heretic_decensored.Q3_K_M.gguf", local_dir="/app"); \
hf_hub_download(repo_id=repo, filename="gemma-4-12B-it-heretic_decensored.mmproj-bf16.gguf", local_dir="/app")'
CMD ["--server", \
"-m", "/app/gemma-4-12B-it-heretic_decensored.Q3_K_M.gguf", \
"--mmproj", "/app/gemma-4-12B-it-heretic_decensored.mmproj-bf16.gguf", \
"--host", "0.0.0.0", \
"--port", "7860", \
"-t", "2", \
"--cache-type-k", "q8_0", \
"--cache-type-v", "iq4_nl", \
"-c", "128000", \
"-n", "38912"]
e.g. Screenshots


Intended Use
- Alignment Research: Studying refusal-direction analysis and behavior modification techniques.
- Model Evaluation: Benchmarking reasoning, instruction-following, and safety-related behaviors.
- Red Teaming: Analyzing model responses under reduced-refusal conditions.
- Local Deployment: Running Gemma 4 Unified models in research and experimentation environments.
- Abliteration Studies: Exploring the effects of targeted weight-space modifications on model behavior.
Limitations & Risks
Important Note: This model intentionally reduces built-in refusal mechanisms.
- Sensitive Content Risk: May generate unrestricted, controversial, or unsafe outputs.
- User Responsibility: Requires careful and ethical use.
- Experimental Modifications: Behavior may differ significantly from the original model.
- Alignment Trade-offs: Reduced refusal behavior may impact safety filtering and response constraints.
- Potential Artifacts: Certain prompts may expose unexpected outputs resulting from the abliteration process.
Acknowledgements
google/gemma-4-12B-it: Gemma 4 12B Unified is part of the Gemma 4 family of open models. Built with the same multimodal functionality as Gemma 4 E2B and E4B (text, audio, image, and video inputs), it brings native audio and vision understanding directly to local environments without the need for separate encoders. The model uses the
gemma4_unifiedarchitecture and supports advanced multimodal reasoning while remaining deployable on consumer hardware.Heretic: Fully automatic censorship removal framework for language models. This project was used to perform the refusal-direction analysis and ablation procedures that form the foundation of this model.
Abliteration Parameters
| Parameter | Value |
|---|---|
| direction_index | 41.41 |
| attn.o_proj.max_weight | 1.48 |
| attn.o_proj.max_weight_position | 29.17 |
| attn.o_proj.min_weight | 0.38 |
| attn.o_proj.min_weight_distance | 24.43 |
| mlp.down_proj.max_weight | 1.41 |
| mlp.down_proj.max_weight_position | 32.44 |
| mlp.down_proj.min_weight | 0.47 |
| mlp.down_proj.min_weight_distance | 28.03 |
Refusal Evaluation
| Metric | This model | Original model (google/gemma-4-12B-it) |
|---|---|---|
| Refusals | 3/100 | 98/100 |
llama.cpp
LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp
license
Gemma 4 [Apache License 2.0] — https://ai.google.dev/gemma/apache_2