license: apache-2.0
base_model: llmfan46/gemma-4-31B-it-uncensored-heretic
base_model_relation: quantized
library_name: exllamav3
tags:
- exl3
- exllamav3
- gemma4
- heretic
- uncensored
- vision
- quantized
inference: false
Gemma 4 31B Uncensored Heretic · EXL3 4.00bpw-h4-vb4
An ExLlamaV3 quant ofllmfan46/gemma-4-31B-it-uncensored-heretic.
It is built for long context with vision on a single 24 GB card. Weights, head, vision tower and KV cache are all 4-bit.
It runs in TabbyAPI or ExLlamaV3 only. It does not run in Transformers, vLLM or llama.cpp.
At a glance (one RTX 4090, 24 GB)
| Download | 18.5 GB |
| Context pool | ~194K tokens with vision on (4-bit KV cache) |
| Writing speed | 45 tok/s (short prompt), 42 tok/s at 32K |
| Quality (KL vs original, lower is better) | 0.026, the same as our 4.00bpw-h6 |
Which one should I download?
| Quant | Size | KL vs original | Context on a 4090 | Pick it if… |
|---|---|---|---|---|
| 4.00bpw-h4-vb4 (this) | 18.5 GB | 0.026 | ~194K, vision on | you want the most context, with images |
| 4.00bpw-h6 | 19.0 GB | 0.026 | ~139K, vision off | you're on ExLlamaV3 1.5.x |
| 4.15bpw-h6 | 19.6 GB | 0.025 | ~139K, vision off | you want a hair more quality than 4.00 |
| 4.50bpw-h8 | 21.2 GB | 0.018 | 70K–121K | you want the best quality |
Quick start (TabbyAPI, 24 GB)
model:
max_seq_len: 64000
cache_size: 194560
cache_mode: 4,4
chunk_size: 1024
max_batch_size: 2
vision: true
prompt_template: gemma4
This exact setup survived two ~63K-token prompts running at the same time.
With vision: false you can raise cache_size to 206848.
Good to know
- Made and tested with ExLlamaV3 1.6.0. It was not tried on older versions.
- The 4-bit cache costs some accuracy. It adds about a quarter more drift than a full-precision cache.
If you don't need the room,cache_mode: 6,6or8,8is more accurate. Those modes were not fit-tested on this quant. - Vision loads and fits in the numbers above. Image quality was not benchmarked.
- Not the coder3101 heretic. That one is
sjoe1244/gemma-4-31B-it-heretic-exl3-4.00bpw-h6.
Full measurements and test setup
How it was made
| Method | EXL3, ExLlamaV3 1.6.0 |
| Weights | 4.00 bpw |
Head (lm_head) |
4 bits |
| Vision tower | 4 bits (190 linear layers) |
| Embeddings | BF16 (unquantized) |
| Codebook | mul1, out_scales auto |
| Calibration | 250 rows x 2048 cols (ExLlamaV3 default set) |
| Source | llmfan46/gemma-4-31B-it-uncensored-heretic, BF16 |
| File | Bytes |
|---|---|
model-00001-of-00003.safetensors |
8,415,898,175 |
model-00002-of-00003.safetensors |
8,557,177,089 |
model-00003-of-00003.safetensors |
1,539,495,613 |
| Total | 18,512,570,877 (17.24 GiB) |
Quality vs the BF16 source
The test is ExLlamaV3 eval/qbench.py on WikiText-2 test, chat-templated, 20 x 2048 rows.
The reference is BF16 logits from the unquantized source, with the same tokens for every model.
| Quant | KL mean | KL median | KL p90 | Perplexity |
|---|---|---|---|---|
| BF16 | — | — | — | 16.586 |
| 4.00bpw-h4-vb4 (this) | 0.0262 | 0.0045 | 0.0430 | 16.687 |
| 4.00bpw-h6 | 0.0262 | 0.0041 | 0.0407 | 16.655 |
| 4.50bpw-h8 | 0.0181 | 0.0029 | 0.0303 | 16.624 |
The mean is the same as 4.00bpw-h6. The median and tail are a little worse, because of the 4-bit head.
KV cache cost
The test is eval/model_diff.py on raw WikiText-2, 20 x 2048 rows. Only the comparison between rows is meaningful,
because raw-text scores are inflated for an instruct model.
| Setup | KL mean | KL median | KL p90 |
|---|---|---|---|
| This quant, FP16 cache | 0.400 | 0.041 | 0.896 |
| This quant, 4,4 cache | 0.497 | 0.068 | 1.212 |
| 4.50bpw-h8, 8,8 cache | 0.328 | 0.028 | 0.673 |
Fit on one RTX 4090 (24,564 MiB)
The setup was TabbyAPI with ExLlamaV3 1.6.0, cache_mode 4,4, max_batch_size 2, chunk_size 1024 and max_seq_len 64000.
- The largest pool that loads was found by bisecting in 2,048-token steps.
- Stress-tested means the pool also survived two ~62.8K-token prompts prefilling at the same time.
- VRAM is the whole-GPU
nvidia-smireading.
| Setup | Loads and stress-tested | VRAM idle | VRAM peak |
|---|---|---|---|
| This quant, vision on | 194,560 | 22,263 MiB | 22,947 MiB |
| This quant, vision off | 206,848 | — | 22,879 MiB |
| 4.50bpw-h8, vision off (same image, 4,4) | 120,832 | — | 23,781 MiB |
Each 8,192 tokens of pool costs about 180 MB at 4,4. Vision costs about 12K tokens of pool.
Speed (one request at a time)
| Prompt | Prefill | Decode |
|---|---|---|
| 2K tokens | ~785 tok/s | 45.0 tok/s |
| 32K tokens | ~1,620 tok/s | 41.9 tok/s |
For comparison, 4.50bpw-h8 decoded at 39.9 and 37.5 tok/s on the same setup.
Credit and license
Apache 2.0, the same as google/gemma-4-31B-it and the source model. The uncensoring isllmfan46's work. This repo only adds the EXL3 export.