← back to catalog · registered 2026-08-22 13:56

culturerevolt/gemma-4-12b-heretic-abliterated-GGUF

culturerevolt Gemma 12B GGUF second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/culturerevolt%2Fgemma-4-12b-heretic-abliterated-GGUF"
Response includes
  • classification m3
  • files 9
  • hub_downloads_all_time 703,615
  • author_summary 10 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
704K
116K last 30d - stable
Likes
32
Model age
4mo ago
created 2026-06-05
Downloads over time
Now738.3K→from5.4K↑13,502%
0270.5K541.1K811.6K5.4K on Jun 10738.3K on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 60 snapshots · spans 123 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 3 formats · 116K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Quantizations
IQ3 IQ4 Q4_K Q5_K Q6_K Q8_0
Tags
gguf gemma-4 abliterated uncensored heretic imatrix multi-quant text-generation base_model:culturerevolt/gemma-4-12b-heretic-abliterated base_model:quantized:culturerevolt/gemma-4-12b-heretic-abliterated license:apache-2.0 endpoints_compatible

Related

Total size
46.8 GB
Files
9
Quantizations
8
Registered
2026-08-22 13:56
Last updated on HF
2026-06-05 11:32

Files by quantization

Q8_0 1 file 11.8 GB
gemma-4-12b-heretic-abliterated-Q8_0.gguf 11.8 GB e4734aeb download
Q6_K 1 file 9.11 GB
gemma-4-12b-heretic-Q6_K.gguf 9.11 GB ea50ec1e download
Q5_K 1 file 7.96 GB
gemma-4-12b-heretic-Q5_K_M.gguf 7.96 GB 02635254 download
Q4_K 1 file 6.87 GB
gemma-4-12b-heretic-Q4_K_M.gguf 6.87 GB 6c4067ea download
IQ4 1 file 6.18 GB
gemma-4-12b-heretic-IQ4_XS.gguf 6.18 GB e0b47213 download
IQ3 1 file 4.91 GB
gemma-4-12b-heretic-IQ3_XS.gguf 4.91 GB e3b98917 download
F16 1 file 167 MB
gemma-4-12b-heretic-mmproj-f16.gguf 167 MB 2e269f90 download
Auxiliary files 2 files 8.63 KB
README.md 6.67 KB acd0340e download
.gitattributes 1.96 KB d6e75058 download

README current version from Hugging Face


license: apache-2.0
base_model: culturerevolt/gemma-4-12b-heretic-abliterated
tags:

  • gemma-4
  • abliterated
  • uncensored
  • heretic
  • gguf
  • imatrix
  • multi-quant
    pipeline_tag: text-generation

Gemma-4-12B-Heretic-Abliterated-iMatrix-GGUF (Complete Suite)

This repository hosts a complete, professional-grade suite of iMatrix (Importance Matrix) GGUF quantizations of culturerevolt/gemma-4-12b-heretic-abliterated.

The base model is an abliterated, fully decensored variant of Google's gemma-4-12b-it architecture, stripped of categorical refusal strings via norm-preserving directional ablation. This comprehensive GGUF suite is calibrated to preserve maximum linguistic entropy, narrative depth for creative fiction, and strict syntax parsing for agentic tools across multiple hardware profiles.


📊 Quick-Reference: Which Quant Should You Download?

All models (except the standard Q8_0 baseline) utilize a custom calculated importance matrix to prevent the loss of reasoning performance that standard static quants suffer from. Use this table to match the model files to your available hardware footprint:

File Name Quant Type Sizing / VRAM Target Best Use Case
gemma-4-12b-heretic-Q8_0.gguf Q8_0 24GB VRAM (RTX 3090 / 4090 / Mac Studio) The Master Reference. Runs at a near-lossless INT8 resolution. Best suited for high-end setups where maximum possible precision is required.
gemma-4-12b-heretic-Q6_K.gguf Q6_K 20GB–24GB VRAM (RTX 3090 / Dual GPUs) The Near-Lossless Sweet Spot. Offers an incredibly high quality retention while saving significant space over the raw weights.
gemma-4-12b-heretic-Q5_K_M.gguf Q5_K_M 16GB+ VRAM (RTX 4080 / 3090 / 4090) The Enthusiast Tier. The highest fidelity compression level before hitting severe diminishing returns. Retains near-flawless parity with the source weights.
gemma-4-12b-heretic-Q4_K_M.gguf Q4_K_M 16GB VRAM (RTX 4080 / 4070 Ti Super) The Classic Standard. Features a balanced distribution of bit weights. Outstanding general purpose reasoning, chat response times, and stability.
gemma-4-12b-heretic-IQ4_XS.gguf IQ4_XS 12GB–16GB VRAM (RTX 4070 / 4080) The Optimizer's Choice. Custom tailored to maximize token throughput while protecting memory headroom. Fits alongside large context windows (up to 131k+) and MCP pipelines without aggressive system RAM spilling.
gemma-4-12b-heretic-IQ3_XS.gguf IQ3_XS 8GB–12GB VRAM / Integrated Graphics / Mobile The Lightweight Champion. Drastically reduces the memory footprint to run on lower-spec hardware, laptops, or mobile backends. The iMatrix lifeline keeps it coherent where flat 3-bit models collapse.

🧠 Technical Methodology: The iMatrix Advantage

When compressing a model down to 3, 4, or 5 bits, a basic, uncalibrated static quantization acts like a global color reduction on an image because it destroys critical structural details. Large Language Models contain sensitive outlier weights that dictate logic patterns, punctuation structure, and instruction adherence. Indiscriminately compressing these weights leads to specialization collapse, which causes a model to loop tokens, output formatting gibberish, or hallucinate heavily.

To protect this repo from degradation, these files were compiled using a Custom Hybrid Importance Matrix (imatrix.dat) strategy.

1. Calibration Profile

The calibration matrix was generated over a high-entropy, multi-domain text dataset specifically balanced for the local LLM stack:

  • Core Logic & Reasoning: Grounded using community standard English linguistic corpuses (text_en_large) to lock down general intelligence, deductive consistency, and vocabulary breadth.
  • Agentic Punctuation: Interwoven with structured tool-calling strings (tools_large) to safeguard brackets, colons, nested variables, and JSON structures necessary for automated file parsing pipelines like VaultForge.
  • Narrative Fidelity: Infused with creative prose samples to explicitly train the importance matrix to prioritize complex character interiority and vivid world-building pacing.

2. Execution Parameters

To avoid calculation errors or PCIe data-transfer distortions, the matrix was evaluated over a tight, isolated memory profile:

  • Total Cycles: 200 high-entropy chunks (representing roughly 100,000 tokens of diverse data)
  • Chunk Boundary: 512 tokens
  • Synchronization Sizing: -b 512 -ub 512 (Forcing physical batch alignment to protect mathematical scaling)

The resulting map acts as an explicit instruction booklet during quantization. It forces the llama-quantize tool to fiercely protect sensitive reasoning neurons while aggressively compressing the easy background language weights.


⚙️ Recommended Inference Settings

The Gemma-4 family is a highly capable but precise architecture. To avoid text stuttering, formatting drops, or interface crashes in backends like LM Studio, AnythingLLM, or llama-server, implement the following configurations:

1. Multi-Modal Vision Execution

Gemma-4 features a unified, encoder-free architecture, projecting raw image patches straight into the embedding layers. To handle vision tasks inside llama.cpp based frontends, you must load the companion multimodal projector file alongside your chosen text quant.

The verified companion file gemma-4-12b-it-mmproj-f16.gguf is hosted right here in the repository root. Ensure you map this file inside your backend Vision Adapter settings slot to seamlessly initialize image ingestion.


📜 Acknowledgements

  • Google DeepMind for pioneering the unified Gemma-4 architecture.
  • Philipp Emanuel Weidmann for developing the underlying Heretic abliteration framework.
  • Massive thanks to the open-source local AI community for continuously pushing the boundaries of what is possible on local consumer hardware.

⚠️ Disclaimer & Boundary Limits

This model is completely unaligned. It will output text without filtering, judgment, or warning labels. By downloading this model, you accept full responsibility for the prompts fed to it and the text generated by it. Use responsibly within local sandbox development setups.


🎛️ Streamlined Jinja Chat Template

If you encounter interface parsing issues with heavy multi-turn configurations or want to maximize token efficiency during rapid back-and-forth chat sessions, use this clean, hyper-efficient template:

{%- for message in messages %}
<|turn|>{{ message['role'] }}
{{ message['content'] }}<|turn|>
{%- endfor %}
{%- if add_generation_prompt %}
<|turn|>assistant<|channel>thought <channel|>
{%- endif %}

README history 7 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-05Update README.mdca1e60b6.7 KB
    Loading...
  2. 2026-06-05Update README.mdf2265ba6.6 KB
    Loading...
  3. 2026-06-05Update README.md6e8a7a76.3 KB
    Loading...
  4. 2026-06-05Update README.mdc76c4736.1 KB
    Loading...
  5. 2026-06-05Update README.md03f821b6.2 KB
    Loading...
  6. 2026-06-05Update README.mdeaa8de96 KB
    Loading...
  7. 2026-06-05Create README.md9463ba51.4 KB
    Loading...

Discussions 6 threads

  1. 2026-10-08Why is it still reject my request? Is it not heretic?open1 💬#6
    Loading...
  2. 2026-09-25it can't answer my question, have not abliteratedopen1 💬#5
    Loading...
  3. 2026-09-24The Q8 model refusing to respondopen1 💬#4
    Loading...
  4. 2026-09-20sorry this just doesn't answer any questionopen1 💬#3
    Loading...
  5. 2026-08-14multimodal support ?open1 💬#2
    Loading...
  6. 2026-08-13Will there be MLX quants of this one?open1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration