license: apache-2.0
base_model: llmfan46/gemma-4-31B-it-uncensored-heretic
base_model_relation: quantized
tags:
- exl3
- exllamav3
- gemma4
- heretic
- uncensored
- vision
- quantized
inference: false
gemma-4-31B-it-uncensored-heretic-exl3-4.15bpw-h6
EXL3 export of llmfan46/gemma-4-31B-it-uncensored-heretic
for ExLlamaV3 / TabbyAPI. Vision tower quantized to 6-bit.
Load with ExLlamaV3 or TabbyAPI's exllamav3 loader. This will not load in
Transformers, vLLM or llama.cpp.
Not the coder3101 heretic — that daily quant lives atsjoe1244/gemma-4-31B-it-heretic-exl3-4.00bpw-h6
(base_model: coder3101/gemma-4-31B-it-heretic). This repo is the uncensored-heretic lineage.
Quantization
| Method | EXL3 |
| ExLlamaV3 | 1.5.3 |
| Weights | 4.15 bpw |
| Head | 6 bits |
| Vision | 6 bpw |
| Codebook | mul1 (default) |
| Calibration | 250 rows x 2048 cols |
Converted on an RTX 4090 on 2026-09-29 from the local BF16 source
(/mnt/ssd2/models/llm/gemma-4-31B-it-uncensored-heretic-bf16, ~59 GB,Gemma4ForConditionalGeneration).
Size
| File | Bytes |
|---|---|
model-00001-of-00003.safetensors |
8,454,131,452 |
model-00002-of-00003.safetensors |
8,512,835,008 |
model-00003-of-00003.safetensors |
2,583,424,211 |
| Total | 19,583,566,237 (18.239 GiB) |
Measured on RTX 4090 (24 GB)
TabbyAPI config: max_seq_len / cache_size 139264, cache_mode 4,4, max_batch_size 2,chunk_size 2048, vision: false (same as the daily gemma31-400-q4.yml lane).
| Metric | Value |
|---|---|
| Peak VRAM @ ~130K prompt | n/a |
| Decode speed (wall, 128 tok) | n/a tok/s |
| WikiText-2 PPL (8×2048, short) | 879.836852 |
| KL vs BF16 (model_diff, 8×2048) | n/a |
Recommended Tabby settings (24 GB)
max_seq_len: 139264
cache_size: 139264
cache_mode: 4,4
chunk_size: 2048
max_batch_size: 2
vision: false
prompt_template: gemma4
If peak VRAM exceeds ~23.5 GB, drop context or use a lower bpw / more aggressive KV.
Credit and license
Apache 2.0, matching google/gemma-4-31B-it and the source model.
Decensoring credit belongs to llmfan46 — this
repo is only the EXL3 export. Quantization by
ExLlamaV3.