← back to catalog · registered 2026-08-22 13:56

DBMe/gemma-4-12B-it-qat-q4_0-unquantized-uncensored-heretic-exl3-mul1

DBMe Gemma 12B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/DBMe%2Fgemma-4-12B-it-qat-q4_0-unquantized-uncensored-heretic-exl3-mul1"
Response includes
  • classification m3
  • files 4
  • hub_downloads_all_time 44
  • author_summary 7 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
44
11 last 30d - stable
Likes
0
Model age
2mo ago
created 2026-07-15
Downloads over time
Now46→from5↑820%
31934505 on Jul 1546 on Oct 1146 on Oct 7JulAugSepOct
Jul 15 → Oct 11 · 53 snapshots · spans 88 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
exllamav3 exl3 quantized text-generation base_model:llmfan46/gemma-4-12B-it-qat-q4_0-unquantized-uncensored-heretic base_model:quantized:llmfan46/gemma-4-12B-it-qat-q4_0-unquantized-uncensored-heretic license:apache-2.0 region:us

Related

Total size
0 B
Files
4
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-16 00:27

Files by quantization

Auxiliary files 4 files 73.3 KB
metrics_graph.png 67.8 KB 16db61e6 download
README.md 3.25 KB b85e83d7 download
.gitattributes 1.48 KB a6344aac download
metrics.json 751 B 53fcb6e2 download

README current version from Hugging Face


base_model: llmfan46/gemma-4-12B-it-qat-q4_0-unquantized-uncensored-heretic
base_model_relation: quantized
quantized_by: DBMe
library_name: exllamav3
pipeline_tag: text-generation
license: apache-2.0
tags:

  • exl3
  • exllamav3
  • quantized
  • text-generation

gemma-4-12B-it-qat-q4_0-unquantized-uncensored-heretic-exl3-mul1

EXL3 (ExLlamaV3) quantizations of llmfan46/gemma-4-12B-it-qat-q4_0-unquantized-uncensored-heretic. All credit for the original model goes to the original authors.

📊 Available Quantizations & VRAM

The model weights are stored in separate branches. Please switch to a branch to download.
Note: VRAM estimates include PyTorch context overhead (~0.8GB) and assume an unquantized FP16 KV cache.

Target BPW Head BPW Branch (Download Link) WikiText-2 PPL (512 ctx)¹ 2K ctx 4K ctx 8K ctx 16K ctx 32K ctx
4.0 h6 4.0bpw_h6 318.1820 ~7.02 GB ~7.05 GB ~7.11 GB ~7.24 GB ~7.49 GB
5.0 h6 5.0bpw_h6 309.7339 ~8.29 GB ~8.32 GB ~8.38 GB ~8.51 GB ~8.76 GB

¹ Evaluated against WikiText-2 with ExLlamaV3 using a strided 512-token context window (-c 512) in llama.cpp parity mode (-g). Lower is better.

(Higher BPW = higher quality, lower BPW = fits in less VRAM).

📥 How to Download

It's recommended to use the huggingface-cli to download specific branches. (Do not use git clone as it will download all branches!)

Ensure you have the CLI installed:

pip install -U "huggingface_hub[cli]"

Download a specific branch (e.g., 5.0bpw_h6):

# Example: Downloading the 5.0bpw_h6 branch
huggingface-cli download DBMe/gemma-4-12B-it-qat-q4_0-unquantized-uncensored-heretic-exl3-mul1 --revision 5.0bpw_h6 --local-dir gemma-4-12B-it-qat-q4_0-unquantized-uncensored-heretic-exl3-mul1-5.0bpw_h6

💻 Supported Engines

These models are highly optimized for modern GPUs and can be run using:

  • TabbyAPI: A fast, OpenAI-compatible API server. (Set model_name to the local folder name you downloaded the branch into)
  • Text-Generation-WebUI: A local web interface. (Select the exllamav3 loader)
  • ExLlamaV3 (Native): Python library for custom integration.

📈 Perplexity Degradation Curve

(Lower is better)
Perplexity Graph

⚙️ Advanced: Quantization Environment & Settings

🔬 Quantization Settings

  • Codebook: mul1

  • Output Scales: always

  • Calibration Rows: 250

  • Calibration Cols: 2048

  • Calibration Dataset: ExLlamaV3 Default (Wiki/C4/Code)

  • High Quality (HQ) Mode: False

  • ExLlamaV3: 1.0.0 (Commit: cb7f2e9)

  • Hardware: NVIDIA A100-SXM4-40GB

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-16Update main README with VRAM Matrix4fa2a073.2 KB
    Loading...
  2. 2026-07-15Update main README with VRAM Matrix2e2a6263 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration