← back to catalog · registered 2026-08-22 13:56

culturerevolt/gemma-4-12b-heretic-abliterated-6bit-mlx

culturerevolt Gemma 12B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/culturerevolt%2Fgemma-4-12b-heretic-abliterated-6bit-mlx"
Response includes
  • classification m3
  • files 12
  • hub_downloads_all_time 1,208
  • author_summary 10 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
1K
140 last 30d - stable
Likes
0
Model age
4mo ago
created 2026-06-05
Downloads over time
Now1.2K→from670↑86%
6418621.1K1.3K670 on Jun 101.2K on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
mlx safetensors gemma4_unified gemma-4 abliterated uncensored heretic text-generation conversational base_model:culturerevolt/gemma-4-12b-heretic-abliterated base_model:quantized:culturerevolt/gemma-4-12b-heretic-abliterated license:apache-2.0

Related

Total size
11.0 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-05 11:20

Files by quantization

Auxiliary files 12 files 11.1 GB
model-00002-of-00003.safetensors 5.00 GB abe0954c download
model-00001-of-00003.safetensors 4.96 GB 4452e0e1 download
model-00003-of-00003.safetensors 1.08 GB 81c683a2 download
tokenizer.json 30.7 MB cc8d3a0c download
model.safetensors.index.json 132 KB ec4ef352 download
config.json 39.0 KB 4a393b81 download
chat_template.jinja 17.1 KB e61bbfe9 download
README.md 2.81 KB 7c8aae4a download
tokenizer_config.json 2.63 KB d31df919 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 868 B 61c53634 download
generation_config.json 255 B 683ff358 download

README current version from Hugging Face


license: apache-2.0
base_model: culturerevolt/gemma-4-12b-heretic-abliterated
tags:

  • gemma-4
  • abliterated
  • uncensored
  • heretic
  • mlx
    pipeline_tag: text-generation

Gemma-4-12B-Heretic-Abliterated-6Bit-MLX

This repository features an Apple Silicon native 6-bit MLX quantization of culturerevolt/gemma-4-12b-heretic-abliterated.

The base model is an abliterated, fully decensored variant of Google's unified multimodal gemma-4-12b-it architecture, stripped of categorical refusal alignments via norm-preserving directional ablation. This configuration is compiled explicitly for high-performance inference on Mac hardware using the MLX framework.


📊 Format & Hardware Recommendations

This file uses uniform group-wise quantization optimized for unified memory architectures. Every tensor is split into precise channel blocks to isolate and protect outlier weights naturally.

  • Quantization Configuration: 6-bit weights compiled with a group size of 64.
  • Target Environment: This choice is optimized for maximum token speed and memory overhead protection on base Apple Silicon chips or devices running tight context windows alongside heavy system tasks.

⚙️ Recommended Inference Settings

To ensure smooth generation, clean formatting, and proper instruction adherence in native Mac backends like LM Studio or the MLX command line tool, implement the following structures:

1. Multi-Modal Native Capabilities

Gemma-4 is built on a unified, encoder-free architecture. It processes visual tokens and audio waveforms natively without requiring an external vision transformer module. Ensure your choice of frontend is configured to look for the companion Gemma-4-12B scale multimodal projector artifacts to fully enable media ingestion capabilities.


📜 Acknowledgements

  • Google DeepMind for pioneering the unified Gemma-4 architecture.
  • Philipp Emanuel Weidmann for developing the underlying Heretic abliteration framework.
  • Massive thanks to the open-source local AI community for continuously pushing the boundaries of what is possible on local consumer hardware.

⚠️ Disclaimer & Boundary Limits

This model is completely unaligned. It will output text without filtering, judgment, or warning labels. By downloading this model, you accept full responsibility for the prompts fed to it and the text generated by it. Use responsibly within local sandbox development setups.


🎛️ Streamlined Jinja Chat Template

If you encounter interface parsing issues with heavy multi-turn configurations, use this clean, hyper-efficient template to ensure maximum token stability:

{%- for message in messages %}
<|turn|>{{ message['role'] }}
{{ message['content'] }}<|turn|>
{%- endfor %}
{%- if add_generation_prompt %}
<|turn|>assistant<|channel>thought <channel|>
{%- endif %}

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-05Update README.md1da54592.8 KB
    Loading...
  2. 2026-06-05Update README.md968fac82.5 KB
    Loading...
  3. 2026-06-05Update README.md45293752.3 KB
    Loading...
  4. 2026-06-05Update README.md49528922.4 KB
    Loading...
  5. 2026-06-05Update README.md27c811b2.4 KB
    Loading...
  6. 2026-06-05Upload folder using huggingface_hub4da490f84 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration