← back to catalog · registered 2026-08-22 13:56

DBMe/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic-exl3-mul1

DBMe Gemma 31B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/DBMe%2Fgemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic-exl3-mul1"
Response includes
  • classification m3
  • files 4
  • hub_downloads_all_time 70
  • author_summary 7 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
70
25 last 30d - stable
Likes
0
Model age
2mo ago
created 2026-07-15
Downloads over time
Now79→from13↑508%
1035608613 on Jul 1579 on Oct 1179 on Oct 10JulAugSepOct
Jul 15 → Oct 11 · 53 snapshots · spans 88 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
exllamav3 exl3 quantized text-generation base_model:llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic base_model:quantized:llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic license:apache-2.0 region:us

Related

Total size
0 B
Files
4
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-15 18:30

Files by quantization

Auxiliary files 4 files 74.9 KB
metrics_graph.png 68.3 KB 93e0f2a1 download
README.md 3.92 KB 9d46df56 download
.gitattributes 1.48 KB a6344aac download
metrics.json 1.21 KB c88a9cd2 download

README current version from Hugging Face


base_model: llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic
base_model_relation: quantized
quantized_by: DBMe
library_name: exllamav3
pipeline_tag: text-generation
license: apache-2.0
tags:

  • exl3
  • exllamav3
  • quantized
  • text-generation

gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic-exl3-mul1

EXL3 (ExLlamaV3) quantizations of llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic. All credit for the original model goes to the original authors.

📊 Available Quantizations & VRAM

The model weights are stored in separate branches. Please switch to a branch to download.
Note: VRAM estimates include PyTorch context overhead (~0.8GB) and assume an unquantized FP16 KV cache.

Target BPW Head BPW Branch (Download Link) WikiText-2 PPL (512 ctx)¹ 2K ctx 4K ctx 8K ctx 16K ctx 32K ctx
3.5 h6 3.5bpw_h6 2335.9614 N/A N/A N/A N/A N/A
3.75 h6 3.75bpw_h6 2312.4667 N/A N/A N/A N/A N/A
4.0 h6 4.0bpw_h6 2097.6864 N/A N/A N/A N/A N/A
5.0 h6 5.0bpw_h6 1995.7812 N/A N/A N/A N/A N/A
6.0 h6 6.0bpw_h6 1995.3135 N/A N/A N/A N/A N/A
8.0 h8 8.0bpw_h8 1985.1562 N/A N/A N/A N/A N/A

¹ Evaluated against WikiText-2 with ExLlamaV3 using a strided 512-token context window (-c 512) in llama.cpp parity mode (-g). Lower is better.

(Higher BPW = higher quality, lower BPW = fits in less VRAM).

📥 How to Download

It's recommended to use the huggingface-cli to download specific branches. (Do not use git clone as it will download all branches!)

Ensure you have the CLI installed:

pip install -U "huggingface_hub[cli]"

Download a specific branch (e.g., 5.0bpw_h6):

# Example: Downloading the 5.0bpw_h6 branch
huggingface-cli download DBMe/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic-exl3-mul1 --revision 5.0bpw_h6 --local-dir gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic-exl3-mul1-5.0bpw_h6

💻 Supported Engines

These models are highly optimized for modern GPUs and can be run using:

  • TabbyAPI: A fast, OpenAI-compatible API server. (Set model_name to the local folder name you downloaded the branch into)
  • Text-Generation-WebUI: A local web interface. (Select the exllamav3 loader)
  • ExLlamaV3 (Native): Python library for custom integration.

📈 Perplexity Degradation Curve

(Lower is better)
Perplexity Graph

⚙️ Advanced: Quantization Environment & Settings

🔬 Quantization Settings

  • Codebook: mul1

  • Output Scales: always

  • Calibration Rows: 250

  • Calibration Cols: 2048

  • Calibration Dataset: ExLlamaV3 Default (Wiki/C4/Code)

  • High Quality (HQ) Mode: False

  • ExLlamaV3: 1.0.0 (Commit: cb7f2e9)

  • Hardware: NVIDIA RTX PRO 6000 Blackwell Server Edition

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-15Update main README with VRAM Matrix1cdec643.9 KB
    Loading...
  2. 2026-07-15Update main README with VRAM Matrix3d8f6243.7 KB
    Loading...
  3. 2026-07-15Update main README with VRAM Matrix9ef2e1d3.6 KB
    Loading...
  4. 2026-07-15Update main README with VRAM Matrix1e8adb23.4 KB
    Loading...
  5. 2026-07-15Update main README with VRAM Matrixe45696b3.2 KB
    Loading...
  6. 2026-07-15Update main README with VRAM Matrixe9b7cd73 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration